diff --git a/docs/flagrelease_en/model_list.txt b/docs/flagrelease_en/model_list.txt index 7eebd885f..e354d5007 100644 --- a/docs/flagrelease_en/model_list.txt +++ b/docs/flagrelease_en/model_list.txt @@ -1,19 +1,22 @@ FlagRelease/AI21-Jamba-1.5-Mini-hygon-FlagOS +FlagRelease/AI21-Jamba-1.5-Mini-metax-FlagOS FlagRelease/AI21-Jamba-1.5-Mini-nvidia-FlagOS FlagRelease/AgentCPM-Explore-hygon-FlagOS -FlagRelease/AgentCPM-Explore-mthreads-FlagOS -FlagRelease/AgentCPM-Report-mthreads-FlagOS -FlagRelease/Apertus-8B-Instruct-2509-mthreads-FlagOS FlagRelease/BAAI-Cardiac-Agent-hygon-FlagOS FlagRelease/Baguettotron-metax-FlagOS +FlagRelease/Baichuan-M2-32B-hygon-FlagOS FlagRelease/C2S-Scale-Gemma-2-27B-hygon-FlagOS FlagRelease/C2S-Scale-Gemma-2-27B-nvidia-FlagOS +FlagRelease/Codestral-22B-v0.1-hygon-FlagOS +FlagRelease/Codestral-22B-v0.1-mthreads-FlagOS +FlagRelease/DeepCoder-14B-Preview-mthreads-FlagOS FlagRelease/DeepSeek-R1-0528-Qwen3-8B-hygon-FlagOS FlagRelease/DeepSeek-R1-0528-Qwen3-8B-metax-FlagOS FlagRelease/DeepSeek-R1-Distill-Llama-8B-hygon-FlagOS FlagRelease/DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS +FlagRelease/DeepSeek-R1-Distill-Qwen-1.5B-metax-FlagOS FlagRelease/DeepSeek-R1-Distill-Qwen-1.5B-nvidia-FlagOS -FlagRelease/DeepSeek-R1-Distill-Qwen-14B-metax-FlagOS +FlagRelease/DeepSeek-R1-Distill-Qwen-14B-mthreads-FlagOS FlagRelease/DeepSeek-R1-Distill-Qwen-32B-FlagOS-Cambricon FlagRelease/DeepSeek-R1-Distill-Qwen-32B-FlagOS-NVIDIA FlagRelease/DeepSeek-R1-FlagOS-Cambricon-BF16 @@ -40,13 +43,16 @@ FlagRelease/DeepSeek-V4-Pro-mthreads-FlagOS FlagRelease/DeepSeek-V4-Pro-nvidia-FlagOS FlagRelease/ERNIE-4.5-0.3B-PT FlagRelease/ERNIE-4.5-0.3B-PT-hygon-FlagOS +FlagRelease/ERNIE-4.5-0.3B-PT-metax-FlagOS FlagRelease/ERNIE-4.5-0.3B-PT-nvidia-FlagOS FlagRelease/ERNIE-4.5-21B-A3B-PT-hygon-FlagOS +FlagRelease/ERNIE-4.5-21B-A3B-PT-metax-FlagOS FlagRelease/ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS FlagRelease/ERNIE-4.5-300B-A47B-PT-FlagOS FlagRelease/Emu3.5-FlagOS -FlagRelease/Fathom-R1-14B-metax-FlagOS +FlagRelease/Fathom-R1-14B-hygon-FlagOS FlagRelease/GLM-4-32B-Base-0414-hygon-FlagOS +FlagRelease/GLM-4-32B-Base-0414-metax-FlagOS FlagRelease/GLM-4-32B-Base-0414-nvidia-FlagOS FlagRelease/GLM-4.5-FlagOS FlagRelease/GLM-5-FP8-FlagOS @@ -84,9 +90,12 @@ FlagRelease/Hy3-zhenwu-FlagOS FlagRelease/Kimi-K2-Instruct-FlagOS FlagRelease/Kimi-K2-Thinking-FlagOS FlagRelease/Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS +FlagRelease/Kimi-Linear-48B-A3B-Instruct-metax-FlagOS FlagRelease/Kimi-Linear-48B-A3B-Instruct-nvidia-FlagOS FlagRelease/LFM2-2.6B-Exp-metax-FlagOS +FlagRelease/MN-12B-Mag-Mell-R1-mthreads-FlagOS FlagRelease/Meta-Llama-3-8B-Instruct-hygon-FlagOS +FlagRelease/Meta-Llama-3-8B-Instruct-metax-FlagOS FlagRelease/Meta-Llama-3-8B-Instruct-nvidia-FlagOS FlagRelease/MiniCPM-V-4-FlagOS FlagRelease/MiniCPM-V-4-metax-FlagOS @@ -96,7 +105,6 @@ FlagRelease/MiniCPM-o-4.5-iluvatar-FlagOS FlagRelease/MiniCPM-o-4.5-metax-FlagOS FlagRelease/MiniCPM-o-4.5-nvidia-FlagOS FlagRelease/MiniCPM-o-4.5-zhenwu-FlagOS -FlagRelease/MiniCPM4-8B-mthreads-FlagOS FlagRelease/MiniCPM5-1B-Armv9-FlagOS FlagRelease/MiniCPM5-1B-ascend-FlagOS FlagRelease/MiniCPM5-1B-hygon-FlagOS @@ -123,11 +131,20 @@ FlagRelease/MiniMax-M3-metax-FlagOS FlagRelease/MiniMax-M3-mthreads-FlagOS FlagRelease/MiniMax-M3-nvidia-FlagOS FlagRelease/MiniMax-M3-zhenwu-FlagOS +FlagRelease/MiroThinker-v1.5-30B-hygon-FlagOS +FlagRelease/MiroThinker-v1.5-30B-metax-FlagOS +FlagRelease/Mistral-Small-24B-Instruct-2501-mthreads-FlagOS FlagRelease/Moonlight-16B-A3B-hygon-FlagOS +FlagRelease/Moonlight-16B-A3B-metax-FlagOS FlagRelease/Moonlight-16B-A3B-nvidia-FlagOS +FlagRelease/Phi-3-medium-128k-instruct-mthreads-FlagOS +FlagRelease/Phi-3-mini-128k-instruct-iluvatar-FlagOS +FlagRelease/Phi-3-mini-4k-instruct-iluvatar-FlagOS FlagRelease/Phi-3.5-MoE-instruct-hygon-FlagOS +FlagRelease/Phi-3.5-MoE-instruct-metax-FlagOS FlagRelease/Phi-3.5-MoE-instruct-nvidia-FlagOS FlagRelease/Phi-3.5-mini-instruct-FlagOS +FlagRelease/Phi-3.5-mini-instruct-iluvatar-FlagOS FlagRelease/Phi-3.5-mini-instruct-metax-FlagOS FlagRelease/QwQ-32B-FlagOS-Cambricon FlagRelease/QwQ-32B-FlagOS-Iluvatar @@ -143,6 +160,8 @@ FlagRelease/Qwen3-235B-A22B-Instruct-2507-FlagOS FlagRelease/Qwen3-235B-A22B-Instruct-2507-hygon-FlagOS FlagRelease/Qwen3-30B-A3B-FlagOS-nvidia FlagRelease/Qwen3-30B-A3B-Iluvatar-FlagOS +FlagRelease/Qwen3-30B-A3B-Instruct-2507-hygon-FlagOS +FlagRelease/Qwen3-30B-A3B-Thinking-2507-hygon-FlagOS FlagRelease/Qwen3-32B-FlagOS FlagRelease/Qwen3-32B-ascend-FlagOS FlagRelease/Qwen3-4B-FlagOS-Ascend @@ -159,6 +178,8 @@ FlagRelease/Qwen3-Next-80B-A3B-Instruct-FlagOS FlagRelease/Qwen3-Next-80B-A3B-Instruct-metax-FlagOS FlagRelease/Qwen3-Omni-30B-A3B-Instruct-FlagOS FlagRelease/Qwen3-VL-235B-A22B-Instruct-FlagOS +FlagRelease/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-hygon-FlagOS +FlagRelease/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-metax-FlagOS FlagRelease/Qwen3.5-35B-A3B-FlagOS FlagRelease/Qwen3.5-35B-A3B-hygon-FlagOS FlagRelease/Qwen3.5-35B-A3B-iluvatar-FlagOS @@ -186,9 +207,12 @@ FlagRelease/RoboBrain2.0-7B-W8A16-FlagOS FlagRelease/RoboBrain2.0-7B-metax-FlagOS FlagRelease/RoboBrain2.5-8B-FlagOS FlagRelease/RoboBrain2.5-8B-ascend-FlagOS +FlagRelease/SOLAR-10.7B-Instruct-v1.0-mthreads-FlagOS FlagRelease/Seed-OSS-36B-Instruct-FlagOS FlagRelease/Seed-OSS-36B-Instruct-hygon-FlagOS +FlagRelease/Seed-OSS-36B-Instruct-metax-FlagOS FlagRelease/Seed-OSS-36B-Instruct-nvidia-FlagOS +FlagRelease/SuperNova-Medius-mthreads-FlagOS FlagRelease/TeleChat3-36B-Thinking-mthreads-FlagOS FlagRelease/ZCK-Qwen3-8B-metax-FlagOS FlagRelease/ZCK-Qwen3-8B-nvidia-FlagOS @@ -196,14 +220,17 @@ FlagRelease/aya-23-8B-hygon-FlagOS FlagRelease/deepseek-r1-1.5b-nvidia-FlagOS FlagRelease/farm_molecular_representation-hygon-FlagOS FlagRelease/farm_molecular_representation-nvidia-FlagOS -FlagRelease/gemma-1.1-7b-it-metax-FlagOS +FlagRelease/gemma-2-27b-it-hygon-FlagOS +FlagRelease/gemma-2-27b-it-metax-FlagOS FlagRelease/gpt-oss-120b-FlagOS +FlagRelease/granite-4.0-micro-iluvatar-FlagOS FlagRelease/grok-2-FlagOS FlagRelease/llama-3-Korean-Bllossom-8B-hygon-FlagOS FlagRelease/materials.smi-ted-hygon-FlagOS FlagRelease/materials.smi-ted-nvidia-FlagOS FlagRelease/phi-4-FlagOS FlagRelease/phi-4-hygon-FlagOS -FlagRelease/phi-4-metax-FlagOS FlagRelease/pi0-FlagOS +FlagRelease/sarvam-m-metax-FlagOS +FlagRelease/sarvam-m-mthreads-FlagOS FlagRelease/step3-FlagOS diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-hygon-FlagOS.md index 5fd415939..15fc847ba 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-hygon-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-hygon-FlagOS.md @@ -40,7 +40,7 @@ Environment Setup ### Download FlagOS Image ```bash -docker pull harbor.baai.ac.cn/external-cooperation/ai21-jamba-1.5-mini-hygon-tree_0.5.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607101038 +docker pull harbor.baai.ac.cn/external-cooperation/ai21-jamba-1.5-mini-hygon-tree_0.5.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607291451 ``` ### Download Open-source Model Weights @@ -61,7 +61,7 @@ docker run \ -v /data/models:/data/models \ -v /opt/hyhal:/opt/hyhal:ro \ -itd \ - harbor.baai.ac.cn/external-cooperation/ai21-jamba-1.5-mini-hygon-tree_0.5.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607101038 \ + harbor.baai.ac.cn/external-cooperation/ai21-jamba-1.5-mini-hygon-tree_0.5.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607291451 \ sleep infinity docker exec -it ai21-jamba-hygon bash diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-metax-FlagOS.md new file mode 100644 index 000000000..5c47c45ee --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-metax-FlagOS.md @@ -0,0 +1,141 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +AI21-Jamba-1.5-Mini is a hybrid SSM-Transformer language model from AI21 Labs. This release provides a FlagOS-enabled deployment package for the MetaX platform, including container image, model weights, runtime cache, operator replacement logs, and benchmark results. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | AI21-Jamba-1.5-Mini-Nvidia-Origin | AI21-Jamba-1.5-Mini-Metax-FlagOS | +|---|---:|---:| +| livebench_new | 27.5% | 26.2% | +| gpqa_generative_cot | 21.8% | 25.1% | +| musr_generative | 29.0% | 29.8% | +| mmlu_pro | 40.1% | 40.8% | +| gpqa_diamond_generative_cot | 17.7% | 19.2% | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/ai21-jamba-mini-muxi-tree_0.5.1-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12-torch_2.8.0-pcp_maca3.3.0.15-gpu_c550-arc_amd64-driver_3.3.12:2606151839 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/AI21-Jamba-1.5-Mini-metax-FlagOS --local_dir /data/AI21-Jamba-1.5-Mini-metax-FlagOS +``` + +### Start the Container +```bash +docker run \ + --name ai21-jamba-mini-muxi-flagos \ + --network=host \ + --privileged=true \ + --shm-size=16g \ + -v /data:/data \ + -v /tmp/rsy:/data/rsy-tmp \ + -itd harbor.baai.ac.cn/external-cooperation/ai21-jamba-mini-muxi-tree_0.5.1-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12-torch_2.8.0-pcp_maca3.3.0.15-gpu_c550-arc_amd64-driver_3.3.12:2606151839 \ + sleep infinity +docker exec -it ai21-jamba-mini-muxi-flagos /bin/bash +``` +### Start the Server +```bash +#建议按实际卡号调整 +MACA_VISIBLE_DEVICES=1,2 \ +TRITON_ALL_BLOCKS_PARALLEL=1 \ +VLLM_PLUGINS=fl \ +export VLLM_FL_FLAGOS_BLACKLIST="mm,mul,masked_fill_,masked_fill" \ +vllm serve \ + --model /data/AI21-Jamba-1.5-Mini-metax-FlagOS \ + --tensor-parallel-size 2 \ + --served-model-name ai21_flagos \ + --port 8131 \ + --enforce-eager + +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8131/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "ai21_flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from AI-ModelScope/AI21-Jamba-1.5-Mini and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Baichuan-M2-32B-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Baichuan-M2-32B-hygon-FlagOS.md new file mode 100644 index 000000000..2b6b55028 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Baichuan-M2-32B-hygon-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Baichuan-M2-32B-hygon-FlagOS-Origin | Baichuan-M2-32B-hygon-FlagOS-FlagOS | +|--------------|-------------------------------------|-------------------------------------| +| GPQA_Diamond | 0 | 68.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.0 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/baichuan-m2-32b-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607311700-v4 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Baichuan-M2-32B-hygon-FlagOS --local_dir /data/Baichuan-M2-32B-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri --group-add video -v /opt/hyhal:/opt/hyhal -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-public/baichuan-m2-32b-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607311700-v4 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/Baichuan-M2-32B-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Baichuan-M2-32B --tensor-parallel-size 2 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Baichuan-M2-32B", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from baichuan-inc/Baichuan-M2-32B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Codestral-22B-v0.1-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Codestral-22B-v0.1-hygon-FlagOS.md new file mode 100644 index 000000000..4e4bcb89f --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Codestral-22B-v0.1-hygon-FlagOS.md @@ -0,0 +1,110 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Codestral-22B-v0.1-hygon-FlagOS-Origin | Codestral-22B-v0.1-hygon-FlagOS-FlagOS | +|--------------|----------------------------------------|----------------------------------------| +| GPQA_Diamond | 0 | 34.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 28.2.2 +28.2.2 +22.04.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/codestral-22b-v0.1-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-none:202607311410-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Codestral-22B-v0.1-hygon-FlagOS --local_dir /data/Codestral-22B-v0.1-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri --group-add video -v /opt/hyhal:/opt/hyhal -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-public/codestral-22b-v0.1-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-none:202607311410-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/Codestral-22B-v0.1-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Codestral-22B-v0.1 --tensor-parallel-size 2 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Codestral-22B-v0.1", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from mistralai/Codestral-22B-v0.1 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Codestral-22B-v0.1-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Codestral-22B-v0.1-mthreads-FlagOS.md new file mode 100644 index 000000000..fd86d6a0b --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Codestral-22B-v0.1-mthreads-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Codestral-22B-v0.1-mthreads-FlagOS-Origin | Codestral-22B-v0.1-mthreads-FlagOS-FlagOS | +|--------------|-------------------------------------------|-------------------------------------------| +| GPQA_Diamond | 0 | 0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.9 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/codestral-22b-v0.1-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.5-server:202608020650-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Codestral-22B-v0.1-mthreads-FlagOS --local_dir /data/Codestral-22B-v0.1-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=16g --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --tmpfs /tmp:exec -e MTHREADS_VISIBLE_DEVICES=all -e MTHREADS_DRIVER_CAPABILITIES=all -v /data:/data harbor.baai.ac.cn/flagrelease-public/codestral-22b-v0.1-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.5-server:202608020650-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/Codestral-22B-v0.1-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Codestral-22B-v0.1 --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Codestral-22B-v0.1", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from mistralai/Codestral-22B-v0.1 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_DeepCoder-14B-Preview-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_DeepCoder-14B-Preview-mthreads-FlagOS.md new file mode 100644 index 000000000..ef8ea27c1 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_DeepCoder-14B-Preview-mthreads-FlagOS.md @@ -0,0 +1,110 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | DeepCoder-14B-Preview-mthreads-FlagOS-Origin | DeepCoder-14B-Preview-mthreads-FlagOS-FlagOS | +|--------------|----------------------------------------------|----------------------------------------------| +| GPQA_Diamond | 0 | 56.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.7 +24.0.7 +22.04.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/deepcoder-14b-preview-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202607310624-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/DeepCoder-14B-Preview-mthreads-FlagOS --local_dir /data/DeepCoder-14B-Preview-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=16g --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --tmpfs /tmp:exec -e MTHREADS_VISIBLE_DEVICES=all -e MTHREADS_DRIVER_CAPABILITIES=all -v /data:/data harbor.baai.ac.cn/flagrelease-public/deepcoder-14b-preview-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202607310624-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/DeepCoder-14B-Preview-FlagOS --host 0.0.0.0 --port 8000 --served-model-name DeepCoder-14B-Preview --tensor-parallel-size 1 --max-model-len 32768 --enforce-eager --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "DeepCoder-14B-Preview", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from agentica-org/DeepCoder-14B-Preview and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS.md index b59e525d1..01b236667 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS.md @@ -71,16 +71,21 @@ docker exec -it DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS bash ``` ### Start the Server ```bash -export HIP_VISIBLE_DEVICES=0 -export VLLM_PLUGINS=fl -export USE_FLAGGEMS=1 -export VLLM_FL_FLAGOS_WHITELIST=softmax,rms_norm,add,sub,gather,masked_fill_,cumsum_out,lt,lt_scalar,where_self,where_self_out,arange_start,zero_,zeros,ones,full,rand_like,index,reciprocal,cos,sin,cat,to_copy,argmax,le,scatter - -vllm serve \ +ulimit -n 2048 && nohup env \ + HIP_VISIBLE_DEVICES=0,1 \ + VLLM_PLUGINS=fl \ + USE_FLAGGEMS=1 \ + VLLM_FL_FLAGOS_WHITELIST="arange_start,lt,where_self_out,argmax,zeros_like,bitwise_or_tensor,scatter,rsub_scalar,ones,cumsum,bitwise_and_tensor,resolve_neg,lt_scalar,sum_dim,add,diff,index,le,masked_fill,where_self,bitwise_not,gather,mul,zero_,nonzero,resolve_conj,cumsum_out,gt_scalar,softmax_out,softmax" \ + vllm serve \ --model /data/models/DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS \ --served-model-name DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS \ + --host 0.0.0.0 \ --port 46840 \ - --enforce-eager + --gpu-memory-utilization 0.90 \ + --trust-remote-code \ + --tensor-parallel-size 2 \ + --enforce-eager \ + > DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS.log 2>&1 & ``` ## Service Invocation diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-metax-FlagOS.md new file mode 100644 index 000000000..9adadb2cd --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-metax-FlagOS.md @@ -0,0 +1,122 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero is trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as an initial stage, and it delivers outstanding reasoning capabilities. Through RL training, DeepSeek-R1-Zero naturally exhibits numerous powerful and intriguing reasoning behaviors. + +Nevertheless, DeepSeek-R1-Zero suffers from issues such as endless repetition, poor readability, and mixed-language outputs. To address these flaws and further boost reasoning performance, we developed DeepSeek-R1, which incorporates cold-start data prior to the RL phase. DeepSeek-R1 achieves performance comparable to OpenAI o1 on mathematical, coding, and reasoning tasks. + +To support the research community, we have open-sourced DeepSeek-R1-Zero, DeepSeek-R1, as well as six dense models distilled from DeepSeek-R1 based on the Llama and Qwen architectures. DeepSeek-R1-Distill-Qwen-32B outperforms OpenAI o1-mini across a wide range of benchmarks, setting a new state-of-the-art record among dense models. +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | DeepSeek-R1-Distill-Qwen-1.5B-Nvidia-Origin | DeepSeek-R1-Distill-Qwen-1.5B-Metax-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| musr_generative | 0.3320 | 0.3373 | +| mmlu_pro | 0.1747 | 0.1734 | +| aime | 0 | 0 | +| livebench_new | 0.1261 | 0.1271 | +| gpqa_generative_cot | 0.0866 | 0.0835 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/deepseek-r1-distill-qwen-1.5b-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2-pcp_maca3.3.0.15-gpu_c550-arc_x86_64-driver_3.3.12:2606251138 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/DeepSeek-R1-Distill-Qwen-1.5B-metax-FlagOS --local_dir /data/DeepSeek-R1-Distill-Qwen-1.5B-metax-FlagOS +``` + +### Start the Container +```bash +docker run -it --name flagos --privileged --net=host --ipc=host -v /data:/data harbor.baai.ac.cn/external-cooperation/deepseek-r1-distill-qwen-1.5b-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2-pcp_maca3.3.0.15-gpu_c550-arc_x86_64-driver_3.3.12:2606251138 bash +docker exec -it flagos bash +``` +### Start the Server +```bash +CUDA_VISIBLE_DEVICES=0 VLLM_PLUGINS=fl USE_FLAGGEMS=1 VLLM_FL_FLAGOS_WHITELIST=max,index,argmax,where_self,where_self_out,gather,lt,lt_scalar,le,scatter,arange_start,ones,full,fill_scalar_,rand_like,zeros,zero_,exponential_,cat,to_copy vllm serve --model /data/DeepSeek-R1-Distill-Qwen-1.5B-metax-FlagOS --served-model-name DeepSeek-R1-Distill-Qwen-1.5B --port 8000 --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "DeepSeek-R1-Distill-Qwen-1.5B", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-14B-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-14B-mthreads-FlagOS.md new file mode 100644 index 000000000..2d0bb0964 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-14B-mthreads-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | DeepSeek-R1-Distill-Qwen-14B-mthreads-FlagOS-Origin | DeepSeek-R1-Distill-Qwen-14B-mthreads-FlagOS-FlagOS | +|--------------|-----------------------------------------------------|-----------------------------------------------------| +| GPQA_Diamond | 0 | 58.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.9 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/deepseek-r1-distill-qwen-14b-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.5-server:202607310557-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/DeepSeek-R1-Distill-Qwen-14B-mthreads-FlagOS --local_dir /data/DeepSeek-R1-Distill-Qwen-14B-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --tmpfs /tmp:exec -e MTHREADS_VISIBLE_DEVICES=all -e MTHREADS_DRIVER_CAPABILITIES=all -v /data:/data harbor.baai.ac.cn/flagrelease-public/deepseek-r1-distill-qwen-14b-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.5-server:202607310557-v2 sleep infinity +``` +### Start the Server +```bash +USE_FLAGGEMS=1 VLLM_PLUGINS=fl vllm serve /data/DeepSeek-R1-Distill-Qwen-14B-FlagOS --host 0.0.0.0 --port 8000 --served-model-name DeepSeek-R1-Distill-Qwen-14B --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code --enforce-eager --reasoning-parser deepseek_r1 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "DeepSeek-R1-Distill-Qwen-14B", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from deepseek-ai/DeepSeek-R1-Distill-Qwen-14B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-metax-FlagOS.md new file mode 100644 index 000000000..bdea45f65 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-metax-FlagOS.md @@ -0,0 +1,139 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +ERNIE-4.5-0.3B-PT is a lightweight foundational LLM from Baidu Wenxin series.This release is fully adapted to Muxi C550 chip with FlagOS acceleration stack.Precompiled Triton cache, FlagGems operator logs and 4k context benchmark data are embedded in the model folder.It supports out-of-the-box vLLM inference via vLLM-plugin-FL, with obvious throughput improvement compared with native implementation on Muxi hardware. +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | ERNIE-4.5-0.3B-PT-Nvidia-Origin | ERNIE-4.5-0.3B-PT-Metax-FlagOS | +|---------------------|---------------------------------|--------------------------------| +| musr_generative | 0.377 | 0.3717 | +| mmlu_pro | 0.1737 | 0.1758 | +| gpqa_generative_cot | 0.25 | 0.2424 | +| livebench_new | 0.1495 | 0.1456 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/ernie-4.5-0.3b-pt-muxi-tree_0.5.1-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12-torch_2.8.0-pcp_maca3.3.0.15-gpu_c550-arc_amd64-driver_3.3.12:2606161504 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/ERNIE-4.5-0.3B-PT-metax-FlagOS --local_dir /data/ERNIE-4.5-0.3B-PT-metax-FlagOS +``` + +### Start the Container +```bash + docker run -d \ + --name ernie-4.5-0.3b-flagos \ + --privileged \ + --net=host \ + --ipc=host \ + -v /usr/local/models:/usr/local/models \ + -v /usr/local/dev:/usr/local/dev \ + -v /data:/data \ + harbor.baai.ac.cn/external-cooperation/ernie-4.5-0.3b-pt-muxi-tree_0.5.1-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12-torch_2.8.0-pcp_maca3.3.0.15-gpu_c550-arc_amd64-driver_3.3.12:2606161504 \ + sleep infinity +docker exec -it ernie-4.5-0.3b-flagos /bin/bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export VLLM_USE_MODELSCOPE=true + +ulimit -n 2048 && nohup env VLLM_FL_FLAGOS_BLACKLIST="masked_fill,masked_fill_,mm,sort,sort_stable,cumsum,cumsum_out,gather,exponential_,arange_start,index" vllm serve /data/ERNIE-4.5-0.3B-PT-metax-FlagOS \ + --served-model-name ernie-4.5-0.3b-flagos \ + --port 9010 \ + --tensor-parallel-size 1 \ + --enforce-eager \ + --trust-remote-code \ + > ernie_flagos.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:9010/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "ernie-4.5-0.3b-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from PaddlePaddle/ERNIE-4.5-0.3B-PT and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-hygon-FlagOS.md index fae79be59..3d49304aa 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-hygon-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-hygon-FlagOS.md @@ -45,7 +45,7 @@ Environment Setup ### Download FlagOS Image ```bash -docker pull harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607021138 +docker pull harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607291650 ``` ### Download Open-source Model Weights @@ -69,7 +69,7 @@ docker run \ --cap-add=SYS_PTRACE \ --security-opt seccomp=unconfined \ -itd \ - harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607021138 \ + harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607291650 \ sleep infinity docker exec -it flagos bash ``` diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-metax-FlagOS.md new file mode 100644 index 000000000..9c8e6a420 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-metax-FlagOS.md @@ -0,0 +1,148 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +The advanced capabilities of the ERNIE 4.5 models, particularly the MoE-based A47B and A3B series, are underpinned by several key technical innovations. + +1. **Multimodal Heterogeneous MoE Pre-Training:** The models are jointly trained on both textual and visual modalities to better capture the nuances of multimodal information and improve performance on tasks involving text understanding and generation, image understanding, and cross-modal reasoning. To achieve this without one modality hindering the learning of another, a *heterogeneous MoE structure* was designed, incorporating *modality-isolated routing*, *router orthogonal loss*, and *multimodal token-balanced loss*. These architectural choices ensure that both modalities are effectively represented, allowing for mutual reinforcement during training. + +2. **Scaling-Efficient Infrastructure:** A novel heterogeneous hybrid parallelism and hierarchical load balancing strategy is introduced for efficient training of ERNIE 4.5 models. By utilizing intra-node expert parallelism, memory-efficient pipeline scheduling, FP8 mixed-precision training, and fine-grained recomputation methods, high pre-training throughput is achieved. For inference, a *multi-expert parallel collaboration* method and a *convolutional code quantization* algorithm are employed to achieve 4-bit/2-bit lossless quantization. Furthermore, PD disaggregation with dynamic role switching is introduced for effective resource utilization, enhancing inference performance for ERNIE 4.5 MoE models. Built on [PaddlePaddle](GitHub - PaddlePaddle/Paddle: PArallel Distributed Deep LEarning: Machine Learning Framework from In), ERNIE 4.5 delivers high-performance inference across a wide range of hardware platforms. + +3. **Modality-Specific Post-Training:** To meet the diverse requirements of real-world applications, variants of the pre-trained model are fine-tuned for specific modalities. The LLMs are optimized for general-purpose language understanding and generation, while the VLMs focus on vision-language understanding and support both thinking and non-thinking modes. Each model employs a combination of *Supervised Fine-tuning (SFT)*, *Direct Preference Optimization (DPO)*, or a modified reinforcement learning method named *Unified Preference Optimization (UPO)* during post-training. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | ERNIE-4.5-21B-A3B-PT-Nvidia-Origin | ERNIE-4.5-21B-A3B-PT-Metax-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| aime | 0.3667 | 0.3333 | +| gpqa_generative_cot | 0.5713 | 0.5763 | +| musr_generative | 0.6296 | 0.6495 | +| mmlu_pro | 0.6644 | 0.6577 | +| livebench_new | 0.5563 | 0.5645 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker Version 28.0.4, build b8034c0 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2-pcp_maca3.3.0.15-gpu_c550-arc_x86_64-driver_3.3.12:2606171358 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/ERNIE-4.5-21B-A3B-PT-metax-FlagOS --local_dir /data/ERNIE-4.5-21B-A3B-PT-metax-FlagOS +``` + +### Start the Container +```bash +docker run -it \ + --device=/dev/dri \ + --device=/dev/mxcd \ + --group-add video \ + --name flagos \ + --device=/dev/mem \ + --network=host \ + --security-opt seccomp=unconfined \ + --security-opt apparmor=unconfined \ + --shm-size '32gb' \ + --ulimit memlock=-1 \ + -v /usr/local/:/usr/local/ \ + -v /data:/data \ + harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2-pcp_maca3.3.0.15-gpu_c550-arc_x86_64-driver_3.3.12:2606171358 +docker exec -it flagos bash +``` +### Start the Server +```bash +nohup env VLLM_FL_FLAGOS_BLACKLIST=masked_fill,masked_fill_,add,mm,sub \ + VLLM_PLUGINS=fl \ + USE_FLAGGEMS=1 \ + vllm serve --model /data/ERNIE-4.5-21B-A3B-PT-metax-FlagOS \ + --tensor-parallel-size 1 \ + --enforce-eager \ + --served-model-name ernie-4.5-21b-a3b-pt-flagos \ + --port 8000 \ + > serve.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "ernie-4.5-21b-a3b-pt-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from PaddlePaddle/ERNIE-4.5-21B-A3B-PT and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS.md index 4d1731083..cddba7174 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS.md @@ -1,11 +1,8 @@ --- -frameworks: -- "" +license: apache-2.0 language: - zh - en -license: apache-2.0 -tasks: [] --- # Introduction The advanced capabilities of the ERNIE 4.5 models, particularly the MoE-based A47B and A3B series, are underpinned by several key technical innovations: @@ -45,7 +42,7 @@ Environment Setup ### Download FlagOS Image ```bash -docker pull harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-nvidia-tree_0.5.0_3.5-gems_5.0.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.158.01:2605111355 +docker pull harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-nvidia-tree_0.5.0_3.5-gems_5.0.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.158.01:2607301330 ``` ### Download Open-source Model Weights @@ -64,8 +61,7 @@ docker run \ --gpus all \ -v /data/models:/data/models \ -itd \ - harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-nvidia-tree_0.5.0_3.5-gems_5.0.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.158.01:2605111355 \ - sleep infinity + harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-nvidia-tree_0.5.0_3.5-gems_5.0.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.158.01:2607301330 sleep infinity docker exec -it flagos bash ``` ### Start the Server @@ -73,7 +69,6 @@ docker exec -it flagos bash export VLLM_PLUGINS=fl export TRITON_ALL_BLOCKS_PARALLEL=1 export USE_FLAGGEMS=1 - vllm serve \ --model /data/models/ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS \ --tensor-parallel-size 1 \ diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Fathom-R1-14B-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Fathom-R1-14B-hygon-FlagOS.md new file mode 100644 index 000000000..605068488 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Fathom-R1-14B-hygon-FlagOS.md @@ -0,0 +1,110 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Fathom-R1-14B-hygon-FlagOS-Origin | Fathom-R1-14B-hygon-FlagOS-FlagOS | +|--------------|-----------------------------------|-----------------------------------| +| GPQA_Diamond | 0 | 64.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 28.2.2 +28.2.2 +22.04.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/fathom-r1-14b-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-none:202607300317-v4 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Fathom-R1-14B-hygon-FlagOS --local_dir /data/Fathom-R1-14B-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri --group-add video -v /opt/hyhal:/opt/hyhal -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-public/fathom-r1-14b-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-none:202607300317-v4 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/Fathom-R1-14B-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Fathom-R1-14B --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code --reasoning-parser deepseek_r1 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Fathom-R1-14B", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from FractalAIResearch/Fathom-R1-14B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-metax-FlagOS.md new file mode 100644 index 000000000..802fa005e --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-metax-FlagOS.md @@ -0,0 +1,159 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforcement learning, we enhanced the model’s performance in instruction following, engineering code, and function calling, thus strengthening the atomic capabilities required for agent tasks. GLM-4-32B-0414 achieves good results in engineering code, Artifact generation, function calling, search-based Q&A, and report generation. In particular, on several benchmarks, such as code generation or specific Q&A tasks, GLM-4-32B-Base-0414 achieves comparable performance with those larger models like GPT-4o and DeepSeek-V3-0324 (671B). +GLM-Z1-32B-0414 is a reasoning model with deep thinking capabilities. This was developed based on GLM-4-32B-0414 through cold start, extended reinforcement learning, and further training on tasks including mathematics, code, and logic. Compared to the base model, GLM-Z1-32B-0414 significantly improves mathematical abilities and the capability to solve complex tasks. During training, we also introduced general reinforcement learning based on pairwise ranking feedback, which enhances the model's general capabilities. +GLM-Z1-Rumination-32B-0414 is a deep reasoning model with rumination capabilities (against OpenAI's Deep Research). Unlike typical deep thinking models, the rumination model is capable of deeper and longer thinking to solve more open-ended and complex problems (e.g., writing a comparative analysis of AI development in two cities and their future development plans). Z1-Rumination is trained through scaling end-to-end reinforcement learning with responses graded by the ground truth answers or rubrics and can make use of search tools during its deep thinking process to handle complex tasks. The model shows significant improvements in research-style writing and complex tasks. +Finally, GLM-Z1-9B-0414 is a surprise. We employed all the aforementioned techniques to train a small model (9B). GLM-Z1-9B-0414 exhibits excellent capabilities in mathematical reasoning and general tasks. Its overall performance is top-ranked among all open-source models of the same size. Especially in resource-constrained scenarios, this model achieves an excellent balance between efficiency and effectiveness, providing a powerful option for users seeking lightweight deployment. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-4-32B-Base-0414-Nvidia-Origin | GLM-4-32B-Base-0414-Metax-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| mmlu | 0.7710 | 0.7710 | +| cmmlu | 0.8325 | 0.832 | +| gsm8k | 0.8719 | 0.8476 | +| leaderboard_bbh | 0.6574 | 0.6549 | +| hellaswag | 0.6469 | 0.6479 | +| truthfulqa_mc1 | 0.3268 | 0.328 | +| winogrande | 0.7908 | 0.7885 | +| commonsense_qa | 0.7715 | 0.7715 | +| piqa | 0.8194 | 0.8172 | +| openbookqa | 0.368 | 0.370 | +| boolq | 0.8801 | 0.8774 | +| arc_easy | 0.8561 | 0.8544 | +| arc_challenge | 0.5922 | 0.5930 | +| minerva_math_algebra | 0.6731 | 0.6824 | +| ceval-valid | 0.8076 | 0.8113 | +| pubmedqa | 0.788 | 0.788 | +| medqa_4options | 0.729 | 0.7258 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker Version 28.0.4, build b8034c0 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/glm-4-32b-base-0414-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2-pcp_maca3.3.0.15-gpu_c550-arc_x86_64-driver_3.3.12:2606180953 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-4-32B-Base-0414-metax-FlagOS --local_dir /data/models/GLM-4-32B-Base-0414-metax-FlagOS +``` + +### Start the Container +```bash +docker run -it \ + --device=/dev/dri \ + --device=/dev/mxcd \ + --group-add video \ + --name flagos \ + --device=/dev/mem \ + --network=host \ + --security-opt seccomp=unconfined \ + --security-opt apparmor=unconfined \ + --shm-size '32gb' \ + --ulimit memlock=-1 \ + -v /usr/local/:/usr/local/ \ + -v /data/models:/data/models \ + harbor.baai.ac.cn/external-cooperation/glm-4-32b-base-0414-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2-pcp_maca3.3.0.15-gpu_c550-arc_x86_64-driver_3.3.12:2606180953 +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +#建议按实际卡号调整 +nohup env VLLM_FL_FLAGOS_BLACKLIST=masked_fill,masked_fill_,mul,to_copy,mm,cat,sub,index_select,where_self_out,reciprocal,lt,where_self,embedding,log_softmax \ + CUDA_VISIBLE_DEVICES=1,2 \ + VLLM_PLUGINS=fl \ + USE_FLAGGEMS=1 \ + vllm serve --model /data/models/GLM-4-32B-Base-0414-metax-FlagOS \ + --tensor-parallel-size 2 \ + --enforce-eager \ + --served-model-name glm-4-32b-base-0414-flagos \ + --port 8000 \ + > serve.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "glm-4-32b-base-0414-flagos", + "prompt": "你好" + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ZhipuAI/GLM-4-32B-Base-0414 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-nvidia-FlagOS.md index 9afc49510..1037a9276 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-nvidia-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-nvidia-FlagOS.md @@ -54,7 +54,7 @@ Environment Setup ### Download FlagOS Image ```bash -docker pull harbor.baai.ac.cn/external-cooperation/glm-4-32b-base-0414-nvidia-tree_0.5.0-gems_0.5.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.133.20:2606091434 +docker pull harbor.baai.ac.cn/external-cooperation/glm-4-32b-base-0414-nvidia-tree_0.5.0-gems_0.5.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.133.20:2607301248 ``` ### Download Open-source Model Weights @@ -73,7 +73,7 @@ docker run -itd \ --privileged=true \ --shm-size=32G \ -v /data/models:/data/models \ - harbor.baai.ac.cn/external-cooperation/glm-4-32b-base-0414-nvidia-tree_0.5.0-gems_0.5.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.133.20:2606091434 \ + harbor.baai.ac.cn/external-cooperation/glm-4-32b-base-0414-nvidia-tree_0.5.0-gems_0.5.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.133.20:2607301248 sleep infinity docker exec -it flagos bash ``` diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS.md index 9e8a79732..c2280e022 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS.md @@ -85,14 +85,12 @@ docker exec -it flagos bash ### Start the Server ```bash -#建议按实际卡号调整 export VLLM_PLUGINS=fl export TRITON_ALL_BLOCKS_PARALLEL=1 export USE_FLAGGEMS=1 export HIP_VISIBLE_DEVICES=2,3 -export VLLM_FL_FLAGOS_WHITELIST="arange_start,lt,where_self_out,argmax,zeros_like,bitwise_or_tensor,scatter,rsub_scalar,ones,cumsum,bitwise_and_tensor,resolve_neg,lt_scalar,sum_dim,add,diff,index,le,masked_fill,where_self,bitwise_not,gather,mul,zero_,nonzero,resolve_conj,cumsum_out,gt_scalar,softmax_out,softmax" -ulimit -n 2048 && nohup vllm serve \ +ulimit -n 2048 && nohup env VLLM_FL_FLAGOS_WHITELIST="arange_start,lt,where_self_out,argmax,zeros_like,bitwise_or_tensor,scatter,rsub_scalar,ones,cumsum,bitwise_and_tensor,resolve_neg,lt_scalar,sum_dim,add,diff,index,le,masked_fill,where_self,bitwise_not,gather,mul,zero_,nonzero,resolve_conj,cumsum_out,gt_scalar,softmax_out,softmax" vllm serve \ --model /data/Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS \ --served-model-name kimi-linear-48b-a3b-instruct-flagos \ --host 0.0.0.0 \ diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-metax-FlagOS.md new file mode 100644 index 000000000..3ceb19eee --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-metax-FlagOS.md @@ -0,0 +1,144 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +Kimi-Linear-48B-A3B-Instruct is a high-efficiency large language model developed by MoonshotAI. Built with an innovative hybrid linear attention architecture and equipped with 48B total parameters, it is specially optimized for long-context comprehension, multi-turn dialogue and complex reasoning scenarios, supporting an ultra-long context window up to 1 million tokens. + +Adopting a 3:1 structural ratio of Kimi Delta Attention and global MLA, this model greatly cuts down KV cache occupancy and improves inference throughput while maintaining strong comprehensive capability. It achieves outstanding results on multiple authoritative benchmarks, natively compatible with Transformers and vLLM frameworks, and can be quickly deployed for long document parsing, knowledge question answering and industrial intelligent conversation services. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Kimi-Linear-48B-A3B-Instruct-Nvidia-Origin | Kimi-Linear-48B-A3B-Instruct-Metax-FlagOS | +| ------------------- | -------------------------------------------------------- | -------------------------------------------------------- | +| aime | 0.4667 | 0.4620 | +| musr_generative | 0.5926 | 0.5542 | +| mmlu_pro | 0.515 | 0.4784 | +| gpqa_generative_cot | 0.4295 | 0.3985 | +| livebench_new | 0.5438 | 0.5231 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/kimi-linear-48b-a3b-instruct-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2_cu128-pcp_cuda12.8-gpu_metax_c550-arc_amd64-driver_3.3.12:2606081508 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Kimi-Linear-48B-A3B-Instruct-metax-FlagOS --local_dir /data/Kimi-Linear-48B-A3B-Instruct-metax-FlagOS +``` + +### Start the Container +```bash +docker run -itd \ + --name=flagos \ + --privileged \ + --network=host \ + -v /data:/data \ + harbor.baai.ac.cn/external-cooperation/kimi-linear-48b-a3b-instruct-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2_cu128-pcp_cuda12.8-gpu_metax_c550-arc_amd64-driver_3.3.12:2606081508 \ + sleep infinity +docker exec -it flagos bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +export CUDA_VISIBLE_DEVICES=0,1 + +export VLLM_FL_FLAGOS_BLACKLIST="sort,mm,mul,masked_fill_" + +ulimit -n 2048 && nohup vllm serve \ +--model /data/Kimi-Linear-48B-A3B-Instruct-metax-FlagOS \ +--served-model-name kimi-linear-48b-a3b-instruct \ +--host 0.0.0.0 \ +--port 8000 \ +--trust-remote-code \ +--tensor-parallel-size 2 \ +--enforce-eager \ +> kimi_flagos.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "kimi-linear-48b-a3b-instruct", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from moonshotai/Kimi-Linear-48B-A3B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_MN-12B-Mag-Mell-R1-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_MN-12B-Mag-Mell-R1-mthreads-FlagOS.md new file mode 100644 index 000000000..e0489b48e --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_MN-12B-Mag-Mell-R1-mthreads-FlagOS.md @@ -0,0 +1,110 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | MN-12B-Mag-Mell-R1-mthreads-FlagOS-Origin | MN-12B-Mag-Mell-R1-mthreads-FlagOS-FlagOS | +|--------------|-------------------------------------------|-------------------------------------------| +| GPQA_Diamond | 0 | 28.57 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.7 +24.0.7 +22.04.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/mn-12b-mag-mell-r1-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202607300814-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/MN-12B-Mag-Mell-R1-mthreads-FlagOS --local_dir /data/MN-12B-Mag-Mell-R1-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=16g --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --tmpfs /tmp:exec -e MTHREADS_VISIBLE_DEVICES=all -e MTHREADS_DRIVER_CAPABILITIES=all -v /data:/data harbor.baai.ac.cn/flagrelease-public/mn-12b-mag-mell-r1-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202607300814-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/MN-12B-Mag-Mell-R1-FlagOS --host 0.0.0.0 --port 8000 --served-model-name MN-12B-Mag-Mell-R1 --tensor-parallel-size 1 --max-model-len 32768 --enforce-eager --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "MN-12B-Mag-Mell-R1", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from inflatebot/MN-12B-Mag-Mell-R1 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-hygon-FlagOS.md index 11fb6faaf..457d9575c 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-hygon-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-hygon-FlagOS.md @@ -73,11 +73,11 @@ docker exec -it llama-3-8b-flagos /bin/bash ``` ### Start the Server ```bash -export VLLM_PLUGINS=fl -export TRITON_ALL_BLOCKS_PARALLEL=1 +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 export USE_FLAGGEMS=1 -export VLLM_FL_FLAGOS_WHITELIST="softmax,rms_norm,add,sub,gather,masked_fill_,cumsum,cumsum_out,lt,lt_scalar,where_self,where_self_out,sum_dim,arange_start,zero_,zeros,ones,full,rand_like,index,reciprocal,cos,sin,cat,to_copy,argmax,le,scatter" -nohup vllm serve /data/models/Meta-Llama-3-8B-Instruct-hygon-FlagOS \ + +ulimit -n 2048 && nohup env VLLM_FL_FLAGOS_WHITELIST="softmax,rms_norm,add,sub,gather,masked_fill_,cumsum,cumsum_out,lt,lt_scalar,where_self,where_self_out,sum_dim,arange_start,zero_,zeros,ones,full,rand_like,index,reciprocal,cos,sin,cat,to_copy,argmax,le,scatter" VLLM_USE_MODELSCOPE=true vllm serve /data/models/Meta-Llama-3-8B-Instruct-hygon-FlagOS \ --served-model-name llama-3-8b-flagos \ --port 8000 \ --max-num-batched-tokens 2048 \ @@ -146,3 +146,4 @@ We warmly welcome global developers to join us: The model weights are derived from LLM-Research/Meta-Llama-3-8B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-metax-FlagOS.md new file mode 100644 index 000000000..978a80a7e --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-metax-FlagOS.md @@ -0,0 +1,134 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +Meta developed and released the Meta Llama 3 family of Large Language Models (LLMs), a suite of generative text models available in pre-trained and instruction-tuned variants with parameter sizes of 8B and 70B. The instruction-tuned Llama 3 models are optimized for dialogue scenarios and outperform many existing open-source chat models on mainstream industry benchmarks. Additionally, great emphasis was placed on enhancing model helpfulness and safety throughout the development process. +Model Developer: Meta +Variants: Llama 3 comes in two parameter sizes (8B and 70B), with both pre-trained and instruction-tuned releases available. +Input: The model only accepts text inputs. +Output: The model generates only text and code. +Model Architecture: Llama 3 is an autoregressive language model built on an optimized Transformer architecture. Its instruction-tuned variants leverage Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) to align outputs with human preferences regarding helpfulness and safety. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Meta-Llama-3-8B-Instruct-Nvidia-Origin | Meta-Llama-3-8B-Instruct-Metax-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| musr_generative | 0.4524 | 0.4392 | +| mmlu_pro | 0.2174 | 0.2207 | +| aime | 0 | 0 | +| livebench_new | 0.2835 | 0.2847 | +| gpqa_generative_cot| 0.3154 | 0.307 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/meta-llama-3-8b-instruct-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2-pcp_maca3.3.0.15-gpu_c550-arc_x86_64-driver_3.3.12:2607280858 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Meta-Llama-3-8B-Instruct-metax-FlagOS --local_dir /data/models/Meta-Llama-3-8B-Instruct-metax-FlagOS +``` + +### Start the Container +```bash +docker run -itd \ + --name=Meta-Llama-3-8B-Instruct-metax-FlagOS \ + --privileged \ + --network=host \ + -v /data:/data \ +harbor.baai.ac.cn/external-cooperation/meta-llama-3-8b-instruct-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2-pcp_maca3.3.0.15-gpu_c550-arc_x86_64-driver_3.3.12:2607280858 \ + sleep infinity + +docker exec -it Meta-Llama-3-8B-Instruct-metax-FlagOS bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +export VLLM_FL_FLAGOS_WHITELIST="ones, zeros, arange_start, true_divide_, cos, sin, reciprocal, true_divide, cumsum_out, copy_, full, lt, scatter, index" +nohup vllm serve /data/models/Meta-Llama-3-8B-Instruct-metax-FlagOS --served-model-name Meta-Llama-3-8B-Instruct-metax-FlagOS --port 8989 --enforce-eager > Llama-3-8B-fl.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8989/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Meta-Llama-3-8B-Instruct-metax-FlagOS", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from LLM-Research/Meta-Llama-3-8B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_MiroThinker-v1.5-30B-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_MiroThinker-v1.5-30B-hygon-FlagOS.md new file mode 100644 index 000000000..e043143aa --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_MiroThinker-v1.5-30B-hygon-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | MiroThinker-v1.5-30B-hygon-FlagOS-Origin | MiroThinker-v1.5-30B-hygon-FlagOS-FlagOS | +|--------------|------------------------------------------|------------------------------------------| +| GPQA_Diamond | 0 | 34.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.0 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/mirothinker-v1.5-30b-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607302018-v4 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/MiroThinker-v1.5-30B-hygon-FlagOS --local_dir /data/MiroThinker-v1.5-30B-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri --group-add video -v /opt/hyhal:/opt/hyhal -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-public/mirothinker-v1.5-30b-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607302018-v4 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/MiroThinker-v1.5-30B-FlagOS --host 0.0.0.0 --port 8000 --served-model-name MiroThinker-v1.5-30B --tensor-parallel-size 2 --max-model-len 32768 --trust-remote-code --reasoning-parser qwen3 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "MiroThinker-v1.5-30B", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from miromind-ai/MiroThinker-v1.5-30B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_MiroThinker-v1.5-30B-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_MiroThinker-v1.5-30B-metax-FlagOS.md new file mode 100644 index 000000000..7e35f5eb9 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_MiroThinker-v1.5-30B-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | MiroThinker-v1.5-30B-metax-FlagOS-Origin | MiroThinker-v1.5-30B-metax-FlagOS-FlagOS | +|--------------|------------------------------------------|------------------------------------------| +| GPQA_Diamond | 0 | 34.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/mirothinker-v1.5-30b-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202608020156-v4 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/MiroThinker-v1.5-30B-metax-FlagOS --local_dir /data/MiroThinker-v1.5-30B-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-public/mirothinker-v1.5-30b-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202608020156-v4 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/MiroThinker-v1.5-30B-FlagOS --host 0.0.0.0 --port 8000 --served-model-name MiroThinker-v1.5-30B --tensor-parallel-size 2 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "MiroThinker-v1.5-30B", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from miromind-ai/MiroThinker-v1.5-30B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Mistral-Small-24B-Instruct-2501-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Mistral-Small-24B-Instruct-2501-mthreads-FlagOS.md new file mode 100644 index 000000000..f6104ec78 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Mistral-Small-24B-Instruct-2501-mthreads-FlagOS.md @@ -0,0 +1,110 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Mistral-Small-24B-Instruct-2501-mthreads-FlagOS-Origin | Mistral-Small-24B-Instruct-2501-mthreads-FlagOS-FlagOS | +|--------------|--------------------------------------------------------|--------------------------------------------------------| +| GPQA_Diamond | 44.0 | 40.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.7 +24.0.7 +22.04.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/mistral-small-24b-instruct-2501-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202608011052-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Mistral-Small-24B-Instruct-2501-mthreads-FlagOS --local_dir /data/Mistral-Small-24B-Instruct-2501-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=16g --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --tmpfs /tmp:exec -e MTHREADS_VISIBLE_DEVICES=all -e MTHREADS_DRIVER_CAPABILITIES=all -v /data:/data harbor.baai.ac.cn/flagrelease-public/mistral-small-24b-instruct-2501-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202608011052-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/Mistral-Small-24B-Instruct-2501-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Mistral-Small-24B-Instruct-2501 --tensor-parallel-size 1 --max-model-len 32768 --enforce-eager --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Mistral-Small-24B-Instruct-2501", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from mistralai/Mistral-Small-24B-Instruct-2501 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-metax-FlagOS.md new file mode 100644 index 000000000..c2791648f --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-metax-FlagOS.md @@ -0,0 +1,182 @@ +--- +language: zh +license: apache-2.0 +library_name: transformers +pipeline_tag: text-generation +tags: +- moonlight +- metax +- vllm +- moe +- FlagOS +--- + +# Introduction + +**Moonlight-16B-A3B** is a high-performance large language model optimized for the MetaX C550 GPU platform. Built on a Mixture of Experts (MoE) architecture, the model features a total of 16 billion parameters with approximately 3 billion active parameters per inference, striking an optimal balance between high performance and efficient inference throughput. Deeply optimized for MetaX GPU hardware, Moonlight-16B-A3B supports the vLLM inference framework, making it well-suited for large-scale deployment scenarios. + +### Integrated Deployment + +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes + +### Consistency Validation + +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public benchmarks. + +# Evaluation Results + +## Benchmark Result + +| Metrics | Moonlight-16B-A3B-Nvidia-Origin | Moonlight-16B-A3B-Metax-FlagOS | +|------------------|--------------------------------|--------------------------------| +| GPQA_Diamond | 0.1384 | 0.1855 | +| LiveBench New | 0.0475 | 0.0626 | +| musr | 0.0159 | 0.0569 | +| mmlu_pro | 0.1986 | 0.3186 | +| aime | 0.0000 | 0.0000 | + +# User Guide + +## Environment Setup + +| Item | Version | +|------|----------| +| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image + +```bash +docker pull harbor.baai.ac.cn/external-cooperation/moonlight-16b-a3b-metax-tree_0.5.1-gems_5.0.2-vllm_0.13.0-plugin_0.1-cx_none-python_3.12.3-torch_2.8.0-cuda12.9-gpu_metax-arc_amd64-driver_3.3.12:202607311100 +``` + +### Download Open-source Model Weights + +```bash +pip install modelscope + +modelscope download \ + --model FlagRelease/Moonlight-16B-A3B-metax-FlagOS \ + --local_dir /data/Moonlight-16B-A3B-metax-FlagOS +``` + +### Start the Container + +```bash +docker run -itd \ + --name=Moonlight-16B-A3B-metax-FlagOS \ + --privileged \ + --network=host \ + -v /data:/data \ + harbor.baai.ac.cn/external-cooperation/moonlight-16b-a3b-metax-tree_0.5.1-gems_5.0.2-vllm_0.13.0-plugin_0.1-cx_none-python_3.12.3-torch_2.8.0-cuda12.9-gpu_metax-arc_amd64-driver_3.3.12:202607311100 \ + sleep infinity + ``` + +### Enter the Container + +```bash +docker exec -it Moonlight-16B-A3B-metax-FlagOS /bin/bash +``` + +### Start the Server +```bash +export VLLM_FL_FLAGOS_BLACKLIST="mm,masked_fill_,masked_fill" +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 + +CUDA_VISIBLE_DEVICES=1,2 MACA_VISIBLE_DEVICES=1,2 \ +nohup vllm serve /data/Moonlight-16B-A3B-metax-FlagOS \ + --served-model-name moonlight-16b-a3b-flagos \ + --port 8003 \ + --trust-remote-code \ + --max-model-len 32768 \ + --generation-config vllm \ + --gpu-memory-utilization 0.7 \ + --tensor-parallel-size 2 \ + --enforce-eager \ + > /workspace/flagos_server.log 2>&1 & +``` + +## Service Invocation + +### Invocation Script + +```bash +curl http://localhost:8003/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "moonlight-16b-a3b-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters: + - API URL: http://localhost:8003/v1 + - Model: moonlight-flagos +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response + +# Technical Overview + +FlagOS is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. + +With core technologies such as FlagScale, together with vllm-plugin-fl, distributed training/inference framework, FlagGems universal operator library, FlagCX communication library, and FlagTree unified compiler, the FlagRelease platform leverages the FlagOS stack to automatically produce and release various combinations of . + +This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. + +## FlagGems + +FlagGems is a high-performance, generic operator library implemented in Triton language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM training and inference across diverse hardware platforms. + +## FlagTree + +FlagTree is an open-source unified compiler for multiple AI chips. It provides unified compilation capabilities across multiple backends and rapidly implements single-repository multi-backend support. + +## FlagScale and vllm-plugin-fl + +FlagScale is a comprehensive toolkit designed to support the entire lifecycle of large models. It integrates capabilities from Megatron-LM and vLLM to provide an end-to-end solution for training and inference. + +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend. + +## FlagCX + +FlagCX is a scalable and adaptive cross-chip communication library for distributed AI workloads. + +## FlagEval Evaluation Framework + +FlagEval is a comprehensive evaluation system and open platform for large models. It supports large-scale benchmark evaluation across NLP, CV, Audio, and Multimodal tasks. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support + +# License + +The model weights are derived from moonshotai/Moonlight-16B-A3B and are open-sourced under the Apache License 2.0:https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3-medium-128k-instruct-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3-medium-128k-instruct-mthreads-FlagOS.md new file mode 100644 index 000000000..6d26d6feb --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3-medium-128k-instruct-mthreads-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Phi-3-medium-128k-instruct-mthreads-FlagOS-Origin | Phi-3-medium-128k-instruct-mthreads-FlagOS-FlagOS | +|--------------|---------------------------------------------------|---------------------------------------------------| +| GPQA_Diamond | 0 | 40.82 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.9 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/phi-3-medium-128k-instruct-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.5-server:202608011004-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Phi-3-medium-128k-instruct-mthreads-FlagOS --local_dir /data/Phi-3-medium-128k-instruct-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=16g --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --tmpfs /tmp:exec -e MTHREADS_VISIBLE_DEVICES=all -e MTHREADS_DRIVER_CAPABILITIES=all -v /data:/data harbor.baai.ac.cn/flagrelease-public/phi-3-medium-128k-instruct-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.5-server:202608011004-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/Phi-3-medium-128k-instruct-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Phi-3-medium-128k-instruct --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Phi-3-medium-128k-instruct", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from microsoft/Phi-3-medium-128k-instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3-mini-128k-instruct-iluvatar-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3-mini-128k-instruct-iluvatar-FlagOS.md new file mode 100644 index 000000000..c20f2bd09 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3-mini-128k-instruct-iluvatar-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Iluvatar** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Phi-3-mini-128k-instruct-iluvatar-FlagOS-Origin | Phi-3-mini-128k-instruct-iluvatar-FlagOS-FlagOS | +|--------------|-------------------------------------------------|-------------------------------------------------| +| GPQA_Diamond | 0 | 38.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 28.1.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-project/phi-3-mini-128k-instruct-iluvatar001-gems5.0.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt210-ixml44-x64-4.5.0:202607292043-v3 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Phi-3-mini-128k-instruct-iluvatar-FlagOS --local_dir /data/Phi-3-mini-128k-instruct-FlagOS +``` + +### Start the Container +```bash +docker run -itd --name flagos --privileged --network=host -v /dev:/dev -v /lib/modules:/lib/modules -v /data:/data harbor.baai.ac.cn/flagrelease-project/phi-3-mini-128k-instruct-iluvatar001-gems5.0.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt210-ixml44-x64-4.5.0:202607292043-v3 +``` +### Start the Server +```bash +vllm serve /data/Phi-3-mini-128k-instruct-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Phi-3-mini-128k-instruct --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Phi-3-mini-128k-instruct", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from microsoft/Phi-3-mini-128k-instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3-mini-4k-instruct-iluvatar-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3-mini-4k-instruct-iluvatar-FlagOS.md new file mode 100644 index 000000000..c91148376 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3-mini-4k-instruct-iluvatar-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Iluvatar** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Phi-3-mini-4k-instruct-iluvatar-FlagOS-Origin | Phi-3-mini-4k-instruct-iluvatar-FlagOS-FlagOS | +|--------------|-----------------------------------------------|-----------------------------------------------| +| GPQA_Diamond | 0 | 40.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 28.1.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-project/phi-3-mini-4k-instruct-iluvatar001-gems5.0.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt210-ixml44-x64-4.5.0:202607300247-v3 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Phi-3-mini-4k-instruct-iluvatar-FlagOS --local_dir /data/Phi-3-mini-4k-instruct-FlagOS +``` + +### Start the Container +```bash +docker run -itd --name flagos --privileged --network=host -v /dev:/dev -v /lib/modules:/lib/modules -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-project/phi-3-mini-4k-instruct-iluvatar001-gems5.0.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt210-ixml44-x64-4.5.0:202607300247-v3 +``` +### Start the Server +```bash +vllm serve /data/Phi-3-mini-4k-instruct-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Phi-3-mini-4k-instruct --tensor-parallel-size 1 --max-model-len 4096 --trust-remote-code --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Phi-3-mini-4k-instruct", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from microsoft/Phi-3-mini-4k-instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-hygon-FlagOS.md index 66b4c6174..f3a46b588 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-hygon-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-hygon-FlagOS.md @@ -58,8 +58,7 @@ docker exec -it phi-3.5-moe-instruct-hygon-flagos /bin/bash export VLLM_PLUGINS=fl && \ export TRITON_ALL_BLOCKS_PARALLEL=1 && \ export USE_FLAGGEMS=1 && \ -export VLLM_FL_FLAGOS_WHITELIST=zero_,zeros,arange,reciprocal,cos,sin,ge_scalar,lt_scalar,bitwise_and_tensor,bitwise_or_tensor,bitwise_not,embedding,index,rand_like,full,argmax,where_self,where_self_out,lt,cumsum_out,le,scatter,to_copy && \ -nohup vllm serve /data/Phi-3.5-MoE-instruct-hygon-FlagOS \ +ulimit -n 2048 && nohup env VLLM_FL_FLAGOS_WHITELIST="zero_,zeros,arange,reciprocal,cos,sin,ge_scalar,lt_scalar,bitwise_and_tensor,bitwise_or_tensor,bitwise_not,embedding,index,rand_like,full,argmax,where_self,where_self_out,lt,cumsum_out,le,scatter,to_copy" vllm serve /data/Phi-3.5-MoE-instruct-hygon-FlagOS \ --served-model-name phi-3.5-moe-instruct-hygon-flagos \ --host 0.0.0.0 \ --port 8230 \ diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-metax-FlagOS.md new file mode 100644 index 000000000..e9775026f --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-metax-FlagOS.md @@ -0,0 +1,139 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +Phi-3.5-MoE is a lightweight, state-of-the-art open model built upon datasets used for Phi-3 - synthetic data and filtered publicly available documents - with a focus on very high-quality, reasoning dense data. The model supports multilingual and comes with 128K context length (in tokens). The model underwent a rigorous enhancement process, incorporating supervised fine-tuning, proximal policy optimization, and direct preference optimization to ensure precise instruction adherence and robust safety measures. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + + +# Evaluation Results +## Benchmark Result +| Metrics | Phi-3.5-MoE-instruct-Nvidia-Origin | Phi-3.5-MoE-instruct-Metax-FlagOS | +|---------------------|--------------------------------|--------------------------------------| +| aime | 0.0334 | 0.0667 | +| gpqa_generative_cot | 0.3171 | 0.3213 | +| mmlu_pro | 0.5336 | 0.5299 | +| musr_generative | 0.5040 | 0.5119 | +| livebench_new | 0.2863 | 0.2769 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 28.0.4 | +| Operating System | Ubuntu 22.04.4 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/phi-3.5-moe-instruct-metax-tree_0.5.1-gems_0.5.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0.devmetax3.3.0.2_cu116_pcp_maca3.3.0.15-gpu_metax3.3.12_arc_x86_64-driver_3.3.12_count8:26061536 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Phi-3.5-MoE-instruct-metax-FlagOS --local_dir /data/Phi-3.5-MoE-instruct-metax-FlagOS +``` + +### Start the Container +```bash +docker run -itd --name=Phi-3.5-MoE-instruct-metax-FlagOS --privileged --network=host -v /data:/data harbor.baai.ac.cn/external-cooperation/phi-3.5-moe-instruct-metax-tree_0.5.1-gems_0.5.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0.devmetax3.3.0.2_cu116_pcp_maca3.3.0.15-gpu_metax3.3.12_arc_x86_64-driver_3.3.12_count8:26061536 sleep infinity +docker exec -it Phi-3.5-MoE-instruct-metax-FlagOS /bin/bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl && \ +export TRITON_ALL_BLOCKS_PARALLEL=1 && \ +export USE_FLAGGEMS=1 && \ +export VLLM_FL_FLAGOS_WHITELIST=arange,argmax,lt,rand_like,zero_,true_divide_,layer_norm,where_self_out,to_copy,cos,exponential_,repeat,true_divide,lt_scalar,pow_scalar,where_self,embedding,bitwise_not,max,bitwise_and_tensor,bitwise_or_tensor,full,cumsum_out,le,scatter,min,index_select,neg,zeros,reciprocal,mean,std,sin && \ +nohup vllm serve /data/Phi-3.5-MoE-instruct-metax-FlagOS \ + --served-model-name phi-3.5-moe-instruct-metax-flagos \ + --host 0.0.0.0 \ + --port 8230 \ + --trust-remote-code \ + --enforce-eager \ + --tensor-parallel-size 2 \ + --gpu-memory-utilization 0.9 \ + --max-model-len 8192 \ + > Phi-3.5-MoE.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8230/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "phi-3.5-moe-instruct-metax-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response + +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support + +# License +The model weights are derived from LLM-Research/Phi-3.5-MoE-instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + + + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-nvidia-FlagOS.md index a07bb1a57..7b25a1cd4 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-nvidia-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-nvidia-FlagOS.md @@ -81,36 +81,21 @@ docker exec -it Phi-3.5-MoE-instruct-nvidia-flagOS /bin/bash ```bash export VLLM_PLUGINS=fl - export TRITON_ALL_BLOCKS_PARALLEL=1 - export USE_FLAGGEMS=1 - export CUDA_VISIBLE_DEVICES=3 -export VLLM_FL_FLAGOS_WHITELIST="min,mean,arange,max,gather,silu_and_mul,moe_sum,moe_align_block_size,softmax,rand_like,where_self_out,where_self,argmax,true_divide_,true_divide,sort,bitwise_not,embedding,cos,sin,std,reciprocal,lt,ge_scalar,abs" - -ulimit -n 2048 && nohup vllm serve \ - ---model /data/models/Phi-3.5-MoE-instruct-nvidia-FlagOS \ - ---served-model-name phi-3.5-moe-instruct-nvidia-flagos \ - ---host 0.0.0.0 \ - ---port 6679 \ - ---max-model-len 10000 \ - ---gpu-memory-utilization 0.95 \ - ---trust-remote-code \ - ---tensor-parallel-size 1 \ - ---enforce-eager \ - -> phi-3.5_flagos.log 2>&1 & +ulimit -n 2048 && nohup env VLLM_FL_FLAGOS_WHITELIST="min,mean,arange,max,gather,silu_and_mul,moe_sum,moe_align_block_size,softmax,rand_like,where_self_out,where_self,argmax,true_divide_,true_divide,sort,bitwise_not,embedding,cos,sin,std,reciprocal,lt,ge_scalar,abs" VLLM_USE_MODELSCOPE=true vllm serve \ + --model /data/models/Phi-3.5-MoE-instruct-nvidia-FlagOS \ + --served-model-name phi-3.5-moe-instruct-nvidia-flagos \ + --host 0.0.0.0 \ + --port 6679 \ + --max-model-len 10000 \ + --gpu-memory-utilization 0.95 \ + --trust-remote-code \ + --tensor-parallel-size 1 \ + --enforce-eager \ + > phi-3.5_flagos.log 2>&1 & ``` diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-iluvatar-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-iluvatar-FlagOS.md new file mode 100644 index 000000000..f865b9077 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-iluvatar-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Iluvatar** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Phi-3.5-mini-instruct-iluvatar-FlagOS-Origin | Phi-3.5-mini-instruct-iluvatar-FlagOS-FlagOS | +|--------------|----------------------------------------------|----------------------------------------------| +| GPQA_Diamond | 0 | 28.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 28.1.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-project/phi-3.5-mini-instruct-iluvatar001-gems5.0.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt210-ixml44-x64-4.5.0:202607301355-v3 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Phi-3.5-mini-instruct-iluvatar-FlagOS --local_dir /data/Phi-3.5-mini-instruct-FlagOS +``` + +### Start the Container +```bash +docker run -itd --name flagos --privileged --network=host -v /dev:/dev -v /lib/modules:/lib/modules -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-project/phi-3.5-mini-instruct-iluvatar001-gems5.0.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt210-ixml44-x64-4.5.0:202607301355-v3 +``` +### Start the Server +```bash +vllm serve /data/Phi-3.5-mini-instruct-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Phi-3.5-mini-instruct --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Phi-3.5-mini-instruct", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from microsoft/Phi-3.5-mini-instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3-30B-A3B-Instruct-2507-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3-30B-A3B-Instruct-2507-hygon-FlagOS.md new file mode 100644 index 000000000..da54c55a9 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3-30B-A3B-Instruct-2507-hygon-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3-30B-A3B-Instruct-2507-hygon-FlagOS-Origin | Qwen3-30B-A3B-Instruct-2507-hygon-FlagOS-FlagOS | +|--------------|-------------------------------------------------|-------------------------------------------------| +| GPQA_Diamond | 0 | 62.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.0 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3-30b-a3b-instruct-2507-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607302131-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3-30B-A3B-Instruct-2507-hygon-FlagOS --local_dir /data/Qwen3-30B-A3B-Instruct-2507-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri --group-add video -v /opt/hyhal:/opt/hyhal -v /data/models/Qwen3-30B-A3B-Instruct-2507:/data/models/Qwen3-30B-A3B-Instruct-2507 -v /data:/data harbor.baai.ac.cn/flagrelease-public/qwen3-30b-a3b-instruct-2507-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607302131-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/Qwen3-30B-A3B-Instruct-2507-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Qwen3-30B-A3B-Instruct-2507 --tensor-parallel-size 2 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Qwen3-30B-A3B-Instruct-2507", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3-30B-A3B-Instruct-2507 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3-30B-A3B-Thinking-2507-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3-30B-A3B-Thinking-2507-hygon-FlagOS.md new file mode 100644 index 000000000..88c9d6b52 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3-30B-A3B-Thinking-2507-hygon-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3-30B-A3B-Thinking-2507-hygon-FlagOS-Origin | Qwen3-30B-A3B-Thinking-2507-hygon-FlagOS-FlagOS | +|--------------|-------------------------------------------------|-------------------------------------------------| +| GPQA_Diamond | 0 | 72.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.0 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3-30b-a3b-thinking-2507-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607310518-v4 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3-30B-A3B-Thinking-2507-hygon-FlagOS --local_dir /data/Qwen3-30B-A3B-Thinking-2507-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri --group-add video -v /opt/hyhal:/opt/hyhal -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-public/qwen3-30b-a3b-thinking-2507-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607310518-v4 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/Qwen3-30B-A3B-Thinking-2507-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Qwen3-30B-A3B-Thinking-2507 --tensor-parallel-size 2 --max-model-len 32768 --trust-remote-code --reasoning-parser qwen3 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Qwen3-30B-A3B-Thinking-2507", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3-30B-A3B-Thinking-2507 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-hygon-FlagOS.md new file mode 100644 index 000000000..cc5690104 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-hygon-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-hygon-FlagOS-Origin | Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-hygon-FlagOS-FlagOS | +|--------------|---------------------------------------------------------------------|---------------------------------------------------------------------| +| GPQA_Diamond | 0 | 78.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.0 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-project/qwen3.5-27b-claude-4.6-opus-reasoning-distilled-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607301035-v3 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-hygon-FlagOS --local_dir /data/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri --group-add video -v /opt/hyhal:/opt/hyhal -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-project/qwen3.5-27b-claude-4.6-opus-reasoning-distilled-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607301035-v3 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled --tensor-parallel-size 2 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-metax-FlagOS.md new file mode 100644 index 000000000..bfcbfab16 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-metax-FlagOS-Origin | Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-metax-FlagOS-FlagOS | +|--------------|---------------------------------------------------------------------|---------------------------------------------------------------------| +| GPQA_Diamond | 0 | 76.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3.5-27b-claude-4.6-opus-reasoning-distilled-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202608011818-v4 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-metax-FlagOS --local_dir /data/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-public/qwen3.5-27b-claude-4.6-opus-reasoning-distilled-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202608011818-v4 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled --tensor-parallel-size 2 --max-model-len 32768 --gpu-memory-utilization 0.9 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.6-27B-metax-FlagOS-Express.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.6-27B-metax-FlagOS-Express.md index f9afe717d..9ff4e6ae7 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.6-27B-metax-FlagOS-Express.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.6-27B-metax-FlagOS-Express.md @@ -76,7 +76,7 @@ export FLAGGEMS_VENDOR=metax export CUDA_VISIBLE_DEVICES=6,7 export VLLM_FL_FLAGOS_WHITELIST=cat,cos,cumsum,fill,full,gather,gt,le,lt,max,mul,sin,softmax,to,where,zeros,zeros_like export VLLM020_CONTIGUOUS_SINGLE_PREFILL=1 -vllm serve /data/models/Qwen3.6-27B/ \ +vllm serve /data/Qwen3.6-27B \ --tensor-parallel-size 2 --port 8000 --trust-remote-code --dtype bfloat16 \ --served-model-name qwen36-27b \ --max-num-batched-tokens 16384 diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_SOLAR-10.7B-Instruct-v1.0-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_SOLAR-10.7B-Instruct-v1.0-mthreads-FlagOS.md new file mode 100644 index 000000000..e367894f4 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_SOLAR-10.7B-Instruct-v1.0-mthreads-FlagOS.md @@ -0,0 +1,110 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | SOLAR-10.7B-Instruct-v1.0-mthreads-FlagOS-Origin | SOLAR-10.7B-Instruct-v1.0-mthreads-FlagOS-FlagOS | +|--------------|--------------------------------------------------|--------------------------------------------------| +| GPQA_Diamond | 0 | 30.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.7 +24.0.7 +22.04.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/solar-10.7b-instruct-v1.0-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202607300209-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/SOLAR-10.7B-Instruct-v1.0-mthreads-FlagOS --local_dir /data/SOLAR-10.7B-Instruct-v1.0-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=16g --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --tmpfs /tmp:exec -e MTHREADS_VISIBLE_DEVICES=all -e MTHREADS_DRIVER_CAPABILITIES=all -v /data:/data harbor.baai.ac.cn/flagrelease-public/solar-10.7b-instruct-v1.0-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202607300209-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/SOLAR-10.7B-Instruct-v1.0-FlagOS --host 0.0.0.0 --port 8000 --served-model-name SOLAR-10.7B-Instruct-v1.0 --tensor-parallel-size 1 --max-model-len 4096 --trust-remote-code --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "SOLAR-10.7B-Instruct-v1.0", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from upstage/SOLAR-10.7B-Instruct-v1.0 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-metax-FlagOS.md new file mode 100644 index 000000000..a6c21d598 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-metax-FlagOS.md @@ -0,0 +1,136 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +Seed-OSS is a series of open-source large language models developed by ByteDance's Seed Team, designed for powerful long-context, reasoning, agent and general capabilities, and versatile developer-friendly features. Although trained with only 12T tokens, Seed-OSS achieves excellent performance on several popular open benchmarks. +We release this series of models to the open-source community under the Apache-2.0 license. +Key Features + Flexible Control of Thinking Budget: Allowing users to flexibly adjust the reasoning length as needed. This capability of dynamically controlling the reasoning length enhances inference efficiency in practical application scenarios. + Enhanced Reasoning Capability: Specifically optimized for reasoning tasks while maintaining balanced and excellent general capabilities. +Agentic Intelligence: Performs exceptionally well in agentic tasks such as tool-using and issue resolving. + Research-Friendly: Given that the inclusion of synthetic instruction data in pre-training may affect the post-training research, we released pre-trained models both with and without instruction data, providing the research community with more diverse options. + Native Long Context: Trained with up-to-512K long context natively. +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Seed-OSS-36B-Instruct-Nvidia-Origin | Seed-OSS-36B-Instruct-Metax-FlagOS | +|---------------------|-------------------------------------|-------------------------------------| +| gpqa_generative_cot | 0.6149 | 0.6007 | +| aime | 0.6667 | 0.6333 | +| musr_generative | 0.5115 | 0.5058 | +| livebench_new | 0.4167 | 0.4008 | +| mmlu_pro | 0.4886 | 0.4803 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/seed-oss-36b-instruct-metax-tree_0.5.1-metax3.0-gems_0.5.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-2.8.0-metax3.3.0.2-pcp_maca3.3.0.15-gpu_metax-arc_amd64-driver_3.3.12:2606101436 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Seed-OSS-36B-Instruct-metax-FlagOS --local_dir /data/Seed-OSS-36B-Instruct-metax-FlagOS +``` + +### Start the Container +```bash +docker run -itd \ + --name=seed-oss-36b-flagos \ + --privileged \ + --network=host \ + -v /data/Seed-OSS-36B-Instruct-metax-FlagOS:/data/Seed-OSS-36B-Instruct-metax-FlagOS \ + harbor.baai.ac.cn/external-cooperation/seed-oss-36b-instruct-metax-tree_0.5.1-metax3.0-gems_0.5.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-2.8.0-metax3.3.0.2-pcp_maca3.3.0.15-gpu_metax-arc_amd64-driver_3.3.12:2606101436 \ + sleep infinity +docker exec -it seed-oss-36b-flagos /bin/bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +export VLLM_FL_FLAGOS_BLACKLIST="sort,masked_fill_,mm,mul,addmm" +nohup vllm serve --model /data/Seed-OSS-36B-Instruct-metax-FlagOS/ --served-model-name seed-oss-36b-flagos --port 8000 --tensor-parallel-size 2 --trust-remote-code --enforce-eager --max-model-len 8192 >eager-seed-oss-gems.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "seed-oss-36b-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from unsloth/Seed-OSS-36B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_SuperNova-Medius-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_SuperNova-Medius-mthreads-FlagOS.md new file mode 100644 index 000000000..5df593d99 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_SuperNova-Medius-mthreads-FlagOS.md @@ -0,0 +1,110 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | SuperNova-Medius-mthreads-FlagOS-Origin | SuperNova-Medius-mthreads-FlagOS-FlagOS | +|--------------|-----------------------------------------|-----------------------------------------| +| GPQA_Diamond | 32.0 | 42.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.7 +24.0.7 +22.04.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/supernova-medius-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202607312206-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/SuperNova-Medius-mthreads-FlagOS --local_dir /data/SuperNova-Medius-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=16g --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --tmpfs /tmp:exec -e MTHREADS_VISIBLE_DEVICES=all -e MTHREADS_DRIVER_CAPABILITIES=all -v /data:/data harbor.baai.ac.cn/flagrelease-public/supernova-medius-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202607312206-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/SuperNova-Medius-FlagOS --host 0.0.0.0 --port 8000 --served-model-name SuperNova-Medius --tensor-parallel-size 1 --max-model-len 32768 --enforce-eager --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "SuperNova-Medius", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from arcee-ai/SuperNova-Medius and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_gemma-2-27b-it-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_gemma-2-27b-it-hygon-FlagOS.md new file mode 100644 index 000000000..9f43d361b --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_gemma-2-27b-it-hygon-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | gemma-2-27b-it-hygon-FlagOS-Origin | gemma-2-27b-it-hygon-FlagOS-FlagOS | +|--------------|------------------------------------|------------------------------------| +| GPQA_Diamond | 0 | 50.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.0 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/gemma-2-27b-it-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607300635-v4 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/gemma-2-27b-it-hygon-FlagOS --local_dir /data/gemma-2-27b-it-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri --group-add video -v /opt/hyhal:/opt/hyhal -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-public/gemma-2-27b-it-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a:202607300635-v4 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/gemma-2-27b-it-FlagOS --host 0.0.0.0 --port 8000 --served-model-name gemma-2-27b-it --tensor-parallel-size 2 --max-model-len 8192 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "gemma-2-27b-it", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from google/gemma-2-27b-it and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_gemma-2-27b-it-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_gemma-2-27b-it-metax-FlagOS.md new file mode 100644 index 000000000..cadf168d9 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_gemma-2-27b-it-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | gemma-2-27b-it-metax-FlagOS-Origin | gemma-2-27b-it-metax-FlagOS-FlagOS | +|--------------|------------------------------------|------------------------------------| +| GPQA_Diamond | 0 | 46.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/gemma-2-27b-it-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202608010938-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/gemma-2-27b-it-metax-FlagOS --local_dir /data/gemma-2-27b-it-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-public/gemma-2-27b-it-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202608010938-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/gemma-2-27b-it-FlagOS --host 0.0.0.0 --port 8000 --served-model-name gemma-2-27b-it --tensor-parallel-size 2 --max-model-len 8192 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "gemma-2-27b-it", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from google/gemma-2-27b-it and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_granite-4.0-micro-iluvatar-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_granite-4.0-micro-iluvatar-FlagOS.md new file mode 100644 index 000000000..62b5acd8e --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_granite-4.0-micro-iluvatar-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Iluvatar** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | granite-4.0-micro-iluvatar-FlagOS-Origin | granite-4.0-micro-iluvatar-FlagOS-FlagOS | +|--------------|------------------------------------------|------------------------------------------| +| GPQA_Diamond | 0 | 28.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 28.1.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-project/granite-4.0-micro-iluvatar001-gems5.0.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt210-ixml44-x64-4.5.0:202607312207-v3 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/granite-4.0-micro-iluvatar-FlagOS --local_dir /data/granite-4.0-micro-FlagOS +``` + +### Start the Container +```bash +docker run -itd --name flagos --privileged --network=host -v /dev:/dev -v /lib/modules:/lib/modules -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-project/granite-4.0-micro-iluvatar001-gems5.0.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt210-ixml44-x64-4.5.0:202607312207-v3 +``` +### Start the Server +```bash +vllm serve /data/granite-4.0-micro-FlagOS --host 0.0.0.0 --port 8000 --served-model-name granite-4.0-micro --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "granite-4.0-micro", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ibm-granite/granite-4.0-micro and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_phi-4-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_phi-4-hygon-FlagOS.md index bfd6badaa..87360c5f8 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_phi-4-hygon-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_phi-4-hygon-FlagOS.md @@ -1,139 +1,65 @@ # Introduction - -**FlagOS** is a unified heterogeneous computing software stack for large models, co-developed with leading global chip manufacturers. With core technologies such as the **FlagScale** distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the FlagOS stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. - -Based on this, the **phi-4-hygon-FlagOS** model is adapted for the Nvidia chip using the FlagOS software stack, enabling: +新模型介绍,待定.... ### Integrated Deployment - -- Deep integration with the open-source [FlagScale framework](https://github.com/FlagOpen/FlagScale) -- Out-of-the-box inference scripts with pre-configured hardware and software parameters -- Released **FlagOS** container image supporting deployment within minutes - +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes ### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. -- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. - -# Technical Overview - -## **FlagScale Distributed Training and Inference Framework** - -FlagScale is an end-to-end framework for large models across heterogeneous computing resources, maximizing computational efficiency and ensuring model validity through core technologies. Its key advantages include: - -- **Unified Deployment Interface:** Standardized command-line tools support one-click service deployment across multiple hardware platforms, significantly reducing adaptation costs in heterogeneous environments. -- **Intelligent Parallel Optimization:** Automatically generates optimal distributed parallel strategies based on chip computing characteristics, achieving dynamic load balancing of computation/communication resources. -- **Seamless Operator Switching:** Deep integration with the FlagGems operator library allows high-performance operators to be invoked via environment variables without modifying model code. - -## **FlagGems Universal Large-Model Operator Library** - -FlagGems is a Triton-based, cross-architecture operator library collaboratively developed with industry partners. Its core strengths include: - -- **Full-stack Coverage**: Over 100 operators, with a broader range of operator types than competing libraries. -- **Ecosystem Compatibility**: Supports 7 accelerator backends. Ongoing optimizations have significantly improved performance. -- **High Efficiency**: Employs unique code generation and runtime optimization techniques for faster secondary development and better runtime performance compared to alternatives. - -## **FlagEval Evaluation Framework** - -FlagEval (Libra)** is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: - - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. - - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. # Evaluation Results - -## Benchmark Result - -| Metrics | phi-4-H100-CUDA | phi-4-hygon-FlagOS | -| ------------------------- | --------------------- | ------------------ | -|AIME-0shot@avg1|0.200|0.200| -|GPQA-0shot@avg1|0.241|0.225| -|MMLU-5shots@avg1|0.713|0.714| -|MUSR-0shot@avg1|0.594|0.574| -|LiveBench-0shot@avg1|0.431|0.422| +## Benchmark Result +| Metrics | phi-4-hygon-FlagOS-Origin | phi-4-hygon-FlagOS-FlagOS | +|--------------|---------------------------|---------------------------| +| GPQA_Diamond | 0 | 70.0 | +| ERQA | - | - | +| Aime24 | - | - | # User Guide +Environment Setup -**Environment Setup** - -| Item | Version | -| ------------- | ------------------------------------------------------------ | -| Docker Version | Docker version 24.0.6, build ed223bc | -| Operating System | Ubuntu 22.04.4 LTS | -| FlagScale | Version: 0.8.0 | -| FlagGems | Version: 3.0 | +| Item | Version | +|------------------|----------------------| +| Docker Version | 28.2.2 +28.2.2 +22.04.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | ## Operation Steps ### Download FlagOS Image - -BE AWARE!, Hygon's FLAGOS image have not decided public-accesible through internet or not. To obtain this image, you can contact us or hygon through issues. ```bash -docker pull harbor.baai.ac.cn/flagrelease-inner/flagrelease-hygon-release-model_phi-4-tree_none-gems_3.0-scale_0.8.0-cx_none-python_3.10.12-torch_2.4.1_das.opt2.dtk2504-pcp_dtk-25.04-gpu_hygon001-arc_amd64-driver_6.3.13-v1.12.0a:2509011036 +docker pull harbor.baai.ac.cn/flagrelease-project/phi-4-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-none:202607301157-v3 ``` ### Download Open-source Model Weights - ```bash pip install modelscope -modelscope download --model LLM-Research/phi-4 --local_dir /share/phi-4 - +modelscope download --model FlagRelease/phi-4-hygon-FlagOS --local_dir /data/phi-4-FlagOS ``` -### Start the inference service - +### Start the Container ```bash -#Container Startup - -docker run -it \ - --name=flagos \ - --network=host \ - --privileged \ - --ipc=host \ - --shm-size=16G \ - --memory="512g" \ - --ulimit stack=-1:-1 \ - --ulimit memlock=-1:-1 \ - --cap-add=SYS_PTRACE \ - --security-opt seccomp=unconfined \ - --device=/dev/kfd \ - --device=/dev/dri \ - --group-add video \ - -u root \ - -v /opt/hyhal:/opt/hyhal \ - -v /share:/share \ - harbor.baai.ac.cn/flagrelease-inner/flagrelease-hygon-release-model_phi-4-tree_none-gems_3.0-scale_0.8.0-cx_none-python_3.10.12-torch_2.4.1_das.opt2.dtk2504-pcp_dtk-25.04-gpu_hygon001-arc_amd64-driver_6.3.13-v1.12.0a:2509011036 \ - /bin/bash +docker run -d --name flagos --net=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri --group-add video -v /opt/hyhal:/opt/hyhal -v /data:/data -v /data:/data harbor.baai.ac.cn/flagrelease-project/phi-4-hygon001-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-none:202607301157-v3 sleep infinity ``` - -### Serve - +### Start the Server ```bash -flagscale serve phi_4 - +VLLM_PLUGINS=fl vllm serve /data/phi-4-FlagOS --host 0.0.0.0 --port 8000 --served-model-name phi-4 --tensor-parallel-size 1 --max-model-len 16384 --trust-remote-code ``` ## Service Invocation - -### API-based Invocation Script - +### Invocation Script ```bash -import openai -openai.api_key = "EMPTY" -openai.base_url = "http://:9010/v1/" -model = "phi-4-hygon-flagos" -messages = [ - {"role": "system", "content": "You are a helpful assistant."}, - {"role": "user", "content": "What's the weather like today?"} -] -response = openai.chat.completions.create( - model=model, - messages=messages, - stream=False, -) -for item in response: - print(item) - +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "phi-4", + "messages": [{"role": "user", "content": "你好"}] + }' ``` + ### AnythingLLM Integration Guide #### 1. Download & Install @@ -152,9 +78,25 @@ for item in response: #### 3. Model Interaction - After model loading is complete: - - Click **"New Conversation"** - - Enter your question (e.g., “Explain the basics of quantum computing”) - - Click the send button to get a response +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. # Contributing @@ -164,8 +106,5 @@ We warmly welcome global developers to join us: 2. Create Pull Requests to contribute code 3. Improve technical documentation 4. Expand hardware adaptation support - # License - -本模型的权重来源于LLM-Research/phi-4,以apache2.0协议https://www.apache.org/licenses/LICENSE-2.0.txt开源。 - +The model weights are derived from microsoft/phi-4 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_sarvam-m-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_sarvam-m-metax-FlagOS.md new file mode 100644 index 000000000..d6cfa34f4 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_sarvam-m-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | sarvam-m-metax-FlagOS-Origin | sarvam-m-metax-FlagOS-FlagOS | +|--------------|------------------------------|------------------------------| +| GPQA_Diamond | 0 | 56.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/sarvam-m-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202608010646-v4 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/sarvam-m-metax-FlagOS --local_dir /data/sarvam-m-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-public/sarvam-m-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202608010646-v4 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/sarvam-m-FlagOS --host 0.0.0.0 --port 8000 --served-model-name sarvam-m --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "sarvam-m", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from sarvamai/sarvam-m and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_sarvam-m-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_sarvam-m-mthreads-FlagOS.md new file mode 100644 index 000000000..df9a58af0 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_sarvam-m-mthreads-FlagOS.md @@ -0,0 +1,110 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | sarvam-m-mthreads-FlagOS-Origin | sarvam-m-mthreads-FlagOS-FlagOS | +|--------------|---------------------------------|---------------------------------| +| GPQA_Diamond | 12.0 | 50.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.7 +24.0.7 +22.04.1 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/sarvam-m-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202608020423-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/sarvam-m-mthreads-FlagOS --local_dir /data/sarvam-m-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=16g --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --tmpfs /tmp:exec -e MTHREADS_VISIBLE_DEVICES=all -e MTHREADS_DRIVER_CAPABILITIES=all -v /data:/data harbor.baai.ac.cn/flagrelease-public/sarvam-m-mthreads001-gems5.3.0-tree0.6.1-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202608020423-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/sarvam-m-FlagOS --host 0.0.0.0 --port 8000 --served-model-name sarvam-m --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "sarvam-m", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from sarvamai/sarvam-m and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt