Installing FlagGems#

1. Prerequisites#

  • You must ensure that the kernel driver and user-space SDK/toolkits for your hardware have been installed and configured properly. This applies to both NVIDIA platforms and other AI accelerator hardware.

  • If you are trying out the integration with vLLM, you will need to install vLLM or its vendor-customized version if any.

You do not need to manually install Python, PyTorch, or Triton. The setup.sh script handles all of these automatically based on the backend you choose.

2. Install from PyPI#

FlagGems can be installed from PyPI using your favorite Python package manager (e.g. pip).

pip install flag_gems

Info

This installs the pure-Python operators from FlagGems.

To use the C++ wrapped operators (which reduce dispatch overhead for performance-critical paths), you can install a prebuilt native extension wheel via an extra:

pip install "flag-gems[cpp-cuda]"

This pulls in the matching flag-gems-cpp-cuda package from the flagOS PyPI index. Replace cpp-cuda with the extension for your vendor: cpp-musa, cpp-npu, cpp-gcu, or cpp-ix.

If a prebuilt wheel is not available for your platform, see Build and install from source.

2.1. Set up backend dependencies with flaggems-setup#

pip install flag_gems installs the pure-Python operator library, but it does not pull in PyTorch or the vendor-specific runtime packages for your accelerator. The flaggems-setup console script — installed alongside FlagGems — completes the environment by installing the PyTorch stack and vendor packages for a chosen backend from the flagOS PyPI index.

This is the recommended follow-up when you installed FlagGems from PyPI (as opposed to installing from source with setup.sh, which already performs these steps for you).

List the available backends:

flaggems-setup --list

Install the dependencies for your backend (the backend keys are the same ones setup.sh accepts):

# NVIDIA CUDA 12.8
flaggems-setup nvidia-cuda128

# Huawei Ascend CANN 9.0.0
flaggems-setup ascend-cann900

# MetaX MACA 3.8.1
flaggems-setup metax-maca3810

Preview the exact commands without running them:

flaggems-setup nvidia-cuda128 --dry-run

By default the script uses uv pip when uv is on your PATH, and falls back to pip otherwise. Override the installer explicitly with --pip:

flaggems-setup nvidia-cuda128 --pip "pip"

flaggems-setup also installs a Triton-family compiler for you. By default it selects FlagTree when the backend provides one, and falls back to Triton otherwise — the same policy as setup.sh. Override the choice with the --compiler flag or the COMPILER environment variable:

# Force vanilla Triton instead of the auto-selected FlagTree
flaggems-setup nvidia-cuda128 --compiler triton
COMPILER=triton flaggems-setup nvidia-cuda128

See Environment variables for the full COMPILER behavior.

3. Build and install from source#

3.1. Clone the source#

git clone https://github.com/flagos-ai/FlagGems
cd FlagGems/

3.2. Run setup.sh#

The setup.sh script is the recommended way to install FlagGems from source. It reads all configuration from src/flag_gems/backends.yaml and automatically:

  • Installs uv (if not present)
  • Installs the correct Python version for your backend
  • Creates a virtual environment (.venv/)
  • Installs build tools, PyTorch, and vendor-specific dependencies
  • Installs FlagGems with the appropriate extras
  • Installs a compiler (FlagTree or Triton)
  • Installs test dependencies
  • Writes backend environment variables into .venv/bin/activate
./setup.sh <backend>

For example:

# NVIDIA CUDA 12.8
./setup.sh nvidia-cuda128

# Huawei Ascend CANN 9.0.0
./setup.sh ascend-cann900

# MetaX MACA 3.8.1
./setup.sh metax-maca3810

To see available backends:

./setup.sh invalid  # prints the list of available backends

After setup completes, activate the environment and start working:

source .venv/bin/activate
pytest tests/test_abs.py -vs

Tips

  • The environment variables for your backend are automatically included in .venv/bin/activate. No separate environment setup step is needed.
  • By default, FlagTree is installed as the compiler when available. To use vanilla Triton instead, set COMPILER=triton before running setup.sh:
    COMPILER=triton ./setup.sh nvidia-cuda128

3.3. Editable install (for development)#

If you are working on the FlagGems project (e.g. developing new operators), you can perform an editable install so that changes to the Python source take effect immediately without reinstalling:

source .venv/bin/activate
uv pip install --no-build-isolation -e .

setup.sh already installs FlagGems in non-editable mode. Run the command above after setup.sh completes if you want to switch to editable mode. The --no-build-isolation flag reuses the build tools already in the venv.

3.4. C++ extensions (optional)#

FlagGems supports C++ wrapped operators for reduced dispatch overhead on performance-critical operations. The C++ extension is a separate per-vendor package (e.g., flag-gems-cpp-cuda) that installs compiled .so files into the flag_gems/ namespace alongside the pure-Python operator implementations.

There are two ways to get the C++ extensions:

Option A: Build from source with setup.sh#

ENABLE_CPP=1 ./setup.sh nvidia-cuda128

setup.sh automatically:

  • Injects the correct vendor name into cpp/pyproject.toml via tools/set_cpp_vendor.sh
  • Sets the appropriate CMAKE_ARGS (-DFLAGGEMS_BACKEND=...)
  • Builds and installs the C++ extension from the cpp/ subdirectory

Option B: Manual build from the cpp/ subdirectory#

The C++ extension uses scikit-build-core as its build-backend and requires CMake, a C++ toolchain, and your vendor's SDK. Build from the cpp/ subdirectory:

# Set the vendor name (cuda, musa, npu, gcu, or ix)
tools/set_cpp_vendor.sh cuda

# Build and install
CMAKE_ARGS="-DFLAGGEMS_BUILD_C_EXTENSIONS=ON -DFLAGGEMS_BACKEND=CUDA" \
  uv pip install --no-build-isolation ./cpp

For manual control over CMake options, see the CMake options reference.

Runtime: enable with USE_C_EXTENSION#

After installation, set the environment variable to activate C++ paths:

export USE_C_EXTENSION=1

Without this, only torch.ops.flag_gems.* and the c_operators pybind module are active; the ATen replacement and flag_gems.enable() C++ branches require it. See the C++ usage guide for details.

4. References#

4.1 Available backends#

The full list of supported backends is defined in src/flag_gems/backends.yaml. Each backend specifies:

  • Python version
  • PyTorch and vendor-specific dependencies
  • Triton / FlagTree compiler packages
  • Runtime environment variables

4.2 Environment variables#

The COMPILER environment variable controls which compiler to use:

ValueBehavior
(unset)Auto: FlagTree if available, otherwise Triton
flagtreeUse FlagTree
tritonUse vendor Triton

The ENABLE_CPP environment variable enables C++ extensions:

ValueBehavior
(unset or 0)Python-only installation (default)
1Build C++ wrapped operators

4.3 CMake options#

When building with C++ extensions (ENABLE_CPP=1), the following CMake options are set automatically by setup.sh. For manual builds, you can pass them via the CMAKE_ARGS environment variable.

OptionDescriptionDefault
FLAGGEMS_BUILD_C_EXTENSIONSBuild C++ extensionsOFF
FLAGGEMS_BACKENDTarget backend (CUDA, IX, MUSA, NPU, GCU)CUDA
FLAGGEMS_BUILD_CTESTSBuild C++ unit testssame as FLAGGEMS_BUILD_C_EXTENSIONS
FLAGGEMS_INSTALLInstall CMake packageON
FLAGGEMS_USE_EXTERNAL_TRITON_JITUse external Triton JIT libraryOFF
FLAGGEMS_USE_EXTERNAL_PYBIND11Use external pybind11ON
FLAGGEMS_BUILD_POINTWISE_DYNAMIC_CPPBuild pointwise dynamic C++ moduleOFF

4.4 scikit-build-core options#

The main flag-gems package uses setuptools as its build-backend. The scikit-build-core tool is used only for the C++ extension built from the cpp/ subdirectory.

The scikit-build-core tool is a build-backend that bridges CMake and the Python build system, making it easier to create Python modules with CMake. Some commonly used environment variables for configuring scikit-build-core include:

  1. SKBUILD_CMAKE_BUILD_TYPE, used to configure the build type of the project. Valid values are Release, Debug, RelWithDebInfo and MinSizeRel;

  2. SKBUILD_BUILD_DIR, which configures the build directory of the project. The default value is build/{cache_tag}, which is defined in pyproject.toml.

Note that for the environment variable SKBUILD_CMAKE_ARGS, multiple options are separated by semicolons (;), whereas for CMAKE_ARGS, they are separated by spaces.

4.5 The libtriton_jit library#

The C++ extension of FlagGems depends on TritonJIT, a library that implements a Triton JIT runtime in C++ and enables calling Triton JIT functions from C++ code.

If you are building with an external TritonJIT, build and install it first, then pass -DTritonJIT_ROOT=<install path> to CMake:

CMAKE_ARGS="-DFLAGGEMS_BUILD_C_EXTENSIONS=ON -DFLAGGEMS_USE_EXTERNAL_TRITON_JIT=ON -DTritonJIT_ROOT=/usr/local/lib/libtriton_jit" \
ENABLE_CPP=1 ./setup.sh nvidia-cuda128