Installing FlagGems#
1. Prerequisites#
You must ensure that the kernel driver and user-space SDK/toolkits for your hardware have been installed and configured properly. This applies to both NVIDIA platforms and other AI accelerator hardware.
If you are trying out the integration with vLLM, you will need to install vLLM or its vendor-customized version if any.
You do not need to manually install Python, PyTorch, or Triton. The
setup.shscript handles all of these automatically based on the backend you choose.
2. Install from PyPI#
FlagGems can be installed from PyPI
using your favorite Python package manager (e.g. pip).
pip install flag_gemsInfo
This installs the pure-Python operators from FlagGems.
To use the C++ wrapped operators (which reduce dispatch overhead for performance-critical paths), you can install a prebuilt native extension wheel via an extra:
pip install "flag-gems[cpp-cuda]"This pulls in the matching
flag-gems-cpp-cudapackage from the flagOS PyPI index. Replacecpp-cudawith the extension for your vendor:cpp-musa,cpp-npu,cpp-gcu, orcpp-ix.If a prebuilt wheel is not available for your platform, see Build and install from source.
2.1. Set up backend dependencies with flaggems-setup#
pip install flag_gems installs the pure-Python operator library, but it does
not pull in PyTorch or the vendor-specific runtime packages for your
accelerator. The flaggems-setup console script — installed alongside FlagGems
— completes the environment by installing the PyTorch stack and vendor packages
for a chosen backend from the flagOS PyPI index.
This is the recommended follow-up when you installed FlagGems from PyPI (as
opposed to installing from source with setup.sh, which already performs these
steps for you).
List the available backends:
flaggems-setup --listInstall the dependencies for your backend (the backend keys are the same ones
setup.sh accepts):
# NVIDIA CUDA 12.8
flaggems-setup nvidia-cuda128
# Huawei Ascend CANN 9.0.0
flaggems-setup ascend-cann900
# MetaX MACA 3.8.1
flaggems-setup metax-maca3810Preview the exact commands without running them:
flaggems-setup nvidia-cuda128 --dry-runBy default the script uses uv pip when uv is on your PATH, and falls back
to pip otherwise. Override the installer explicitly with --pip:
flaggems-setup nvidia-cuda128 --pip "pip"flaggems-setup also installs a Triton-family compiler for you. By default it
selects FlagTree when the backend provides one, and falls back to Triton
otherwise — the same policy as setup.sh. Override the choice with the
--compiler flag or the COMPILER environment variable:
# Force vanilla Triton instead of the auto-selected FlagTree
flaggems-setup nvidia-cuda128 --compiler triton
COMPILER=triton flaggems-setup nvidia-cuda128See Environment variables for the full COMPILER behavior.
3. Build and install from source#
3.1. Clone the source#
git clone https://github.com/flagos-ai/FlagGems
cd FlagGems/3.2. Run setup.sh#
The setup.sh script is the recommended way to install FlagGems from source.
It reads all configuration from src/flag_gems/backends.yaml and automatically:
- Installs uv (if not present)
- Installs the correct Python version for your backend
- Creates a virtual environment (
.venv/) - Installs build tools, PyTorch, and vendor-specific dependencies
- Installs FlagGems with the appropriate extras
- Installs a compiler (FlagTree or Triton)
- Installs test dependencies
- Writes backend environment variables into
.venv/bin/activate
./setup.sh <backend>For example:
# NVIDIA CUDA 12.8
./setup.sh nvidia-cuda128
# Huawei Ascend CANN 9.0.0
./setup.sh ascend-cann900
# MetaX MACA 3.8.1
./setup.sh metax-maca3810To see available backends:
./setup.sh invalid # prints the list of available backendsAfter setup completes, activate the environment and start working:
source .venv/bin/activate
pytest tests/test_abs.py -vsTips
- The environment variables for your backend are automatically included in
.venv/bin/activate. No separate environment setup step is needed.- By default, FlagTree is installed as the compiler when available. To use vanilla Triton instead, set
COMPILER=tritonbefore running setup.sh:COMPILER=triton ./setup.sh nvidia-cuda128
3.3. Editable install (for development)#
If you are working on the FlagGems project (e.g. developing new operators), you can perform an editable install so that changes to the Python source take effect immediately without reinstalling:
source .venv/bin/activate
uv pip install --no-build-isolation -e .
setup.shalready installs FlagGems in non-editable mode. Run the command above aftersetup.shcompletes if you want to switch to editable mode. The--no-build-isolationflag reuses the build tools already in the venv.
3.4. C++ extensions (optional)#
FlagGems supports C++ wrapped operators for reduced dispatch overhead on
performance-critical operations. The C++ extension is a separate per-vendor
package (e.g., flag-gems-cpp-cuda) that installs compiled .so files into
the flag_gems/ namespace alongside the pure-Python operator implementations.
There are two ways to get the C++ extensions:
Option A: Build from source with setup.sh#
ENABLE_CPP=1 ./setup.sh nvidia-cuda128setup.sh automatically:
- Injects the correct vendor name into
cpp/pyproject.tomlviatools/set_cpp_vendor.sh - Sets the appropriate
CMAKE_ARGS(-DFLAGGEMS_BACKEND=...) - Builds and installs the C++ extension from the
cpp/subdirectory
Option B: Manual build from the cpp/ subdirectory#
The C++ extension uses scikit-build-core as its build-backend and requires
CMake, a C++ toolchain, and your vendor's SDK. Build from the cpp/
subdirectory:
# Set the vendor name (cuda, musa, npu, gcu, or ix)
tools/set_cpp_vendor.sh cuda
# Build and install
CMAKE_ARGS="-DFLAGGEMS_BUILD_C_EXTENSIONS=ON -DFLAGGEMS_BACKEND=CUDA" \
uv pip install --no-build-isolation ./cppFor manual control over CMake options, see the CMake options reference.
Runtime: enable with USE_C_EXTENSION#
After installation, set the environment variable to activate C++ paths:
export USE_C_EXTENSION=1Without this, only torch.ops.flag_gems.* and the c_operators pybind module
are active; the ATen replacement and flag_gems.enable() C++ branches require
it. See the C++ usage guide for details.
4. References#
4.1 Available backends#
The full list of supported backends is defined in src/flag_gems/backends.yaml.
Each backend specifies:
- Python version
- PyTorch and vendor-specific dependencies
- Triton / FlagTree compiler packages
- Runtime environment variables
4.2 Environment variables#
The COMPILER environment variable controls which compiler to use:
| Value | Behavior |
|---|---|
| (unset) | Auto: FlagTree if available, otherwise Triton |
flagtree | Use FlagTree |
triton | Use vendor Triton |
The ENABLE_CPP environment variable enables C++ extensions:
| Value | Behavior |
|---|---|
| (unset or 0) | Python-only installation (default) |
1 | Build C++ wrapped operators |
4.3 CMake options#
When building with C++ extensions (ENABLE_CPP=1), the following CMake
options are set automatically by setup.sh. For manual builds, you can
pass them via the CMAKE_ARGS environment variable.
| Option | Description | Default |
|---|---|---|
FLAGGEMS_BUILD_C_EXTENSIONS | Build C++ extensions | OFF |
FLAGGEMS_BACKEND | Target backend (CUDA, IX, MUSA, NPU, GCU) | CUDA |
FLAGGEMS_BUILD_CTESTS | Build C++ unit tests | same as FLAGGEMS_BUILD_C_EXTENSIONS |
FLAGGEMS_INSTALL | Install CMake package | ON |
FLAGGEMS_USE_EXTERNAL_TRITON_JIT | Use external Triton JIT library | OFF |
FLAGGEMS_USE_EXTERNAL_PYBIND11 | Use external pybind11 | ON |
FLAGGEMS_BUILD_POINTWISE_DYNAMIC_CPP | Build pointwise dynamic C++ module | OFF |
4.4 scikit-build-core options#
The main
flag-gemspackage usessetuptoolsas its build-backend. Thescikit-build-coretool is used only for the C++ extension built from thecpp/subdirectory.
The scikit-build-core tool is a build-backend that bridges CMake
and the Python build system, making it easier to create Python modules with CMake.
Some commonly used environment variables for configuring scikit-build-core include:
SKBUILD_CMAKE_BUILD_TYPE, used to configure the build type of the project. Valid values areRelease,Debug,RelWithDebInfoandMinSizeRel;SKBUILD_BUILD_DIR, which configures the build directory of the project. The default value isbuild/{cache_tag}, which is defined inpyproject.toml.
Note that for the environment variable SKBUILD_CMAKE_ARGS, multiple options
are separated by semicolons (;), whereas for CMAKE_ARGS, they are separated
by spaces.
4.5 The libtriton_jit library#
The C++ extension of FlagGems depends on TritonJIT, a library that implements a Triton JIT runtime in C++ and enables calling Triton JIT functions from C++ code.
If you are building with an external TritonJIT, build and install it first,
then pass -DTritonJIT_ROOT=<install path> to CMake:
CMAKE_ARGS="-DFLAGGEMS_BUILD_C_EXTENSIONS=ON -DFLAGGEMS_USE_EXTERNAL_TRITON_JIT=ON -DTritonJIT_ROOT=/usr/local/lib/libtriton_jit" \
ENABLE_CPP=1 ./setup.sh nvidia-cuda128