Skip to content

Installation

DNALLM is a comprehensive, open-source toolkit designed for fine-tuning and inference with DNA Language Models. This guide will help you install DNALLM and its dependencies.

Prerequisites

  • Python 3.10 or higher (Python 3.13 recommended)
  • Git
  • CUDA-compatible GPU (optional, for GPU acceleration)
  • Environment Manager: Choose one of the following:
  • Python venv (built-in)
  • Conda/Miniconda (recommended for scientific computing)

DNALLM uses uv for dependency management and packaging.

What is uv is a fast Python package manager that is 10-100x faster than traditional tools like pip.

Method 1: Using venv + uv

# Clone repository
git clone https://github.com/zhangtaolab/DNALLM.git
cd DNALLM

# Create virtual environment
python -m venv .venv

# Activate virtual environment
source .venv/bin/activate  # Linux/MacOS
# or
.venv\Scripts\activate     # Windows

# Upgrade pip (recommended)
pip install --upgrade pip

# Install uv in virtual environment
pip install uv

# Install DNALLM with base dependencies
uv pip install -e '.[base]'

# Verify installation
python -c "import dnallm; print('DNALLM installed successfully!')"

Method 2: Using conda + uv

# Clone repository
git clone https://github.com/zhangtaolab/DNALLM.git
cd DNALLM

# Create conda environment
conda create -n dnallm python=3.13 -y

# Activate conda environment
conda activate dnallm

# Install uv in conda environment
conda install uv -c conda-forge

# Install DNALLM with base dependencies
uv pip install -e '.[base]'

# Verify installation
python -c "import dnallm; print('DNALLM installed successfully!')"

Method 3: Using conda + pip

# Clone repository
git clone https://github.com/zhangtaolab/DNALLM.git
cd DNALLM

# Create conda environment
conda create -n dnallm python=3.13 -y

# Activate conda environment
conda activate dnallm

pip install dnallm

# Verify installation
python -c "import dnallm; print('DNALLM installed successfully!')"

Important for GPU users: plain pip does NOT read the per-CUDA-version package indexes configured in pyproject.toml ([tool.uv.sources]), which are only honored by uv. With pip you must install the GPU build of PyTorch explicitly from the PyTorch index — see the GPU Support section below. Otherwise pip silently installs the default (CPU) wheel and torch.cuda.is_available() stays False even with a GPU present.

GPU Support

For GPU acceleration, install the appropriate CUDA version:

# For venv users: activate virtual environment
source .venv/bin/activate  # Linux/MacOS
# or
.venv\Scripts\activate     # Windows

# For conda users: activate conda environment
# conda activate dnallm

# CUDA 12.4 (recommended for recent GPUs)
uv pip install -e '.[cuda124]'

# CUDA 13.0 (requires torch >= 2.9, NVIDIA driver >= 580)
uv pip install -e '.[cuda130]'

# Other supported versions: cpu, cuda121, cuda126, cuda128
uv pip install -e '.[cuda121]'

GPU Support with plain pip

If you use pip instead of uv (e.g., inside a conda environment), the [tool.uv.sources] index configuration is ignored, so the hardware extras above will not pull the GPU build of PyTorch. Install torch explicitly from the PyTorch wheel index instead:

# NVIDIA CUDA (choose one index matching your GPU/driver)
pip install torch --index-url https://download.pytorch.org/whl/cu130

# Intel GPU (XPU, e.g. Intel Arc / Data Center GPU)
pip install torch --index-url https://download.pytorch.org/whl/xpu

# CPU only
pip install torch --index-url https://download.pytorch.org/whl/cpu

Two rules to avoid the most common pitfalls:

  1. Install torch LAST. Run pip install -e '.[base]' (or pip install dnallm) FIRST, then install torch from the index above. The editable install re-resolves torch from PyPI, whose default Windows wheel is the CPU build (+cpu) — it will silently replace a GPU torch installed earlier. Respect the torch<2.12 pin from pyproject.toml, e.g.:
pip install -e '.[base]'                                  # dnallm + deps
pip install --force-reinstall --no-deps 'torch==2.11.0' \
    --index-url https://download.pytorch.org/whl/cu130    # GPU torch LAST
  1. Always verify afterwards — the suffix matters:
python -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.is_available())"
# CUDA env:  2.11.0+cu130 13.0 True
# XPU env:   2.11.0+xpu None False  (check torch.xpu.is_available() instead)
# Wrong:     2.11.0+cpu  None False  <-- pip fell back to the CPU wheel

Note: The torch==2.11.0 pin above is the newest release satisfying DNALLM's torch<2.12 requirement. The PyTorch indexes also ship newer versions (2.12+), which must NOT be used with the current DNALLM release. Check the exact pins in pyproject.toml for your version.

Windows with CUDA 13.0

CUDA 13.0 (cu130) PyTorch wheels are available for Windows and Linux starting from PyTorch 2.9.0:

# 1. Install the latest NVIDIA driver (>= 580.xx) from https://www.nvidia.com/drivers
#    No local CUDA toolkit is required — PyTorch wheels bundle the CUDA runtime.

# 2. Create and activate a virtual environment
python -m venv .venv
.venv\Scripts\activate

# 3. Install DNALLM with CUDA 13.0 support
pip install uv
uv pip install -e '.[base,cuda130]'

# 4. Verify GPU is detected
python -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.is_available())"

Notes for Windows users: - CUDA 13.0 wheels require an NVIDIA driver version 580 or newer. Check with nvidia-smi. - If you have an RTX 50-series (Blackwell) GPU, both cuda128 (torch 2.6+) and cuda130 (torch 2.9+) work; cuda130 ships the newest CUDA runtime. - The local CUDA toolkit (nvcc) is NOT needed for normal usage — only for compiling packages from source (e.g., flash-attn, mamba-ssm), which additionally requires Visual Studio Build Tools (Desktop development with C++).

Dependency Groups

DNALLM provides multiple dependency groups for different use cases:

Core Dependency Groups

Group Purpose Includes
all Install all optional dependencies base + dev + test + notebook + docs + ui + mcp
base Full development environment dev + test + notebook + mcp + extra tools (isort, types-transformers)
dev Complete development environment test + notebook + linting/typing (ruff, flake8, pre-commit, mypy, pandas-stubs)
test Testing environment only pytest and plugins
notebook Jupyter and Marimo support Jupyter Lab, Marimo
docs Documentation building mkdocs-material, mkdocstrings, mkdocs-jupyter
ui Gradio web interface Gradio
mcp MCP server support Included in core dependencies (no extra install needed)

Note: mcp is an empty extra because MCP dependencies (mcp, starlette, uvicorn, websockets) are already part of the core dependencies. You can still use .[mcp] for clarity but it won't install additional packages.

Hardware-Specific Groups (Mutually Exclusive)

Warning: These groups are mutually exclusive. You MUST choose exactly one. Combining multiple hardware groups will cause conflicts.

Group PyTorch Version GPU Type When to Use
cpu 2.4.0-2.7 CPU only Development without GPU
cuda121 2.2.0-2.7 NVIDIA (older) Volta/Turing/Ampere early
cuda124 2.4.0-2.7 NVIDIA (recommended) Most modern GPUs
cuda126 2.6.0-2.7 NVIDIA (latest) Ada/Hopper with Flash Attention
cuda128 2.6.0-2.7 NVIDIA (cutting-edge) RTX 5090 and latest hardware
cuda130 2.9.0-2.12 NVIDIA (CUDA 13.0, Windows & Linux) Newest driver / RTX 50-series, Windows with driver >= 580
rocm 2.5.0-2.7 AMD GPUs AMD GPU users
mamba 2.6.0-2.7 NVIDIA + Mamba Native Mamba architecture (requires CUDA)

Note: Hardware groups are NOT included in all because they conflict with each other. Always combine a hardware group with your chosen feature group: e.g., .[all,cuda124]

Installation Scenarios

Scenario 1: CPU-only Development

For development and testing without GPU acceleration:

# Create environment
conda create -n dnallm-cpu python=3.13 uv -y
conda activate dnallm-cpu

# Install all dependencies and CPU version
uv pip install -e '.[all,cpu]'

# Verify installation
python -c "import dnallm; print('DNALLM installed successfully!')"

Scenario 2: Using NVIDIA GPU for Training and Inference

For GPU-accelerated training and inference:

# Determine CUDA version
nvidia-smi

# Create environment (using CUDA 12.4 as example)
conda create -n dnallm-gpu python=3.13 uv -y
conda activate dnallm-gpu

# Install all dependencies and CUDA 12.4 support
uv pip install -e '.[all,cuda124]'

# Verify installation
python -c "import torch; print(f'PyTorch: {torch.__version__}'); print(f'CUDA available: {torch.cuda.is_available()}')"

Scenario 3: Using Intel GPU (XPU) for Training and Inference

For Intel Arc / Data Center GPU accelerated training and inference:

# Create environment
conda create -n dnallm-xpu python=3.13 -y
conda activate dnallm-xpu

# Install base dependencies first
pip install jupyterlab -e '.[base]'

# Install the XPU build of PyTorch LAST (plain pip only —
# uv users can simply do: uv pip install -e '.[base]' then the xpu index)
pip install --force-reinstall --no-deps 'torch==2.11.0' \
    --index-url https://download.pytorch.org/whl/xpu

# Verify installation
python -c "
import torch
print(f'PyTorch: {torch.__version__}')
print(f'XPU available: {torch.xpu.is_available()}')
if torch.xpu.is_available():
    print(f'GPU: {torch.xpu.get_device_name(0)}')
"

Notes for Intel GPU users: - During training, Hugging Face Trainer automatically falls back to XPU when no CUDA device is present. To force XPU on a machine that ALSO has an NVIDIA GPU, hide CUDA before starting Python: export CUDA_VISIBLE_DEVICES="" (Linux) or set CUDA_VISIBLE_DEVICES= (Windows cmd). - The XPU wheel pulls in Intel SYCL/oneAPI runtime packages automatically. If you ever force-reinstall a different torch version over an existing XPU install, the runtime versions may drift and torch fails to load with OSError: [WinError 126] ... c10_xpu.dll (or shm.dll). Fix by reinstalling the exact runtime pins listed in the torch wheel's METADATA (e.g., intel-sycl-rt==2025.3.2, tbb==2022.3.1, ...), or simply reinstall torch WITH its dependencies (omit --no-deps). - Windows: the Intel GPU driver must be recent enough for the installed PyTorch XPU build — update from Intel's website if torch.xpu.is_available() returns False.

Scenario 4: Multiple GPU Types on One Machine

On a machine with BOTH an NVIDIA GPU and an Intel GPU, the CUDA and XPU torch builds cannot coexist in one environment. Create separate environments and select the kernel per notebook:

# NVIDIA environment
conda create -n dnallm-cuda python=3.13 -y
conda activate dnallm-cuda
uv pip install -e '.[all,cuda130]'        # or plain pip: see Scenario 2 / GPU Support with plain pip

# Intel environment
conda create -n dnallm-xpu python=3.13 -y
conda activate dnallm-xpu
pip install jupyterlab -e '.[base]'
pip install 'torch==2.11.0' --index-url https://download.pytorch.org/whl/xpu

# CPU fallback environment
conda create -n dnallm-cpu python=3.13 -y
conda activate dnallm-cpu
uv pip install -e '.[all,cpu]'

Then pick the matching kernel (dnallm-cuda / dnallm-xpu / dnallm-cpu) in Jupyter/VS Code. The example notebooks include a device-selection cell that auto-detects which devices are visible in the active environment.

Scenario 5: Using Huawei Ascend NPU for Training and Inference

For Huawei Ascend NPU accelerated training and inference, users should first check their device and environment, then install the appropriate dependencies.

If Huawei Ascend driver is not installed in the machine, first check the device and install the corresponding drivers.

For NPU driver, please refer to: https://www.hiascend.com/hardware/firmware-drivers/community

For Ascend Extension for PyTorch (CANN driver), please refer to: https://www.hiascend.com/zh/cann/download

For example, if you have a Ascend 910B NPU with AArch64 architecture, install drivers like this:

# Install NPU driver
wget -c "https://ascend-repo.obs.cn-east-2.myhuaweicloud.com/Ascend%20HDK/Ascend%20HDK%2025.5.2/Ascend-hdk-910b-npu-driver_25.5.2_linux-aarch64.run"
bash ./Ascend-hdk-910b-npu-driver_25.5.2_linux-aarch64.run --install
# Install CANN Toolkit
wget https://ascend-cann-open.obs.cn-north-4.myhuaweicloud.com/CANN/CANN%209.0.0/Ascend-cann_9.0.0_linux-aarch64.run
bash ./Ascend-cann_9.0.0_linux-aarch64.run --install
# Install CANN Kernel/Ops
wget https://ascend-repo.obs.cn-east-2.myhuaweicloud.com/CANN/CANN%209.0.0/Ascend-cann-910b-ops_9.0.0_linux-aarch64.run
bash ./Ascend-cann-910b-ops_9.0.0_linux-aarch64.run --install

# Check the drivers
source /usr/local/Ascend/cann/set_env.sh
python3 -c "import acl;print(acl.get_soc_name())"
npu-smi info

if you want to auto-activate the driver environment, add the set_env.sh to system environment.

echo "source /usr/local/Ascend/cann/set_env.sh" >> ~/.bashrc

To use the NPU accelerating in torch, a specific version of torch_npu package is also required. Please refer to this page to check the dependency map.

For example, CANN 9.0.0 support Pytorch version from 2.7.1 to 2.10.0, also the Python version need to >=3.9.

# Create environment (using CANN 9.0.0 as example)
conda create -n dnallm-npu python=3.11 uv -y
conda activate dnallm-npu

# Install dependencies of NPU support
uv pip install torch torch_npu==2.9.0

# Verify installation
python -c "import torch; import torch_npu; print(f'PyTorch: {torch.__version__}'); print(f'NPU available: {torch.npu.is_available()}')"

During training or inference, Huawei Ascend NPU accelerate is supported for most of the DNA models (models supported by Huggingface Transformers library).

For other non-transformer models or CUDA-dependent models, Huawei provides a specific framework for efficient model training and inference, named MindSpeed. Detailed supported model list is shown here.

Scenario 6: Using Mamba Model Architecture

For models with Mamba architecture (Plant DNAMamba, Caduceus, Jamba-DNA):

# Create environment
conda create -n dnallm-mamba python=3.12 -y
conda activate dnallm-mamba

# Install base dependencies first
uv pip install -e '.[base]'

# Install Mamba support (requires GPU)
uv pip install -e '.[cuda124,mamba]' --no-cache-dir --no-build-isolation

# Verify installation
python -c "from mambapy import Mamba; print('Mamba installed successfully!')"

Scenario 7: Complete Development Environment

For contributors and developers:

# Create environment
conda create -n dnallm-dev python=3.13 -y
conda activate dnallm-dev

# Install all dependencies + CUDA support
uv pip install -e '.[all,cuda124]'

# Verify installation
python -c "
import dnallm
import torch
print('DNALLM:', dnallm.__version__)
print('PyTorch:', torch.__version__)
print('CUDA:', torch.version.cuda if torch.cuda.is_available() else 'CPU')
"

Scenario 8: Running MCP Server Only

For MCP server deployment:

# Create environment
conda create -n dnallm-mcp python=3.13 -y
conda activate dnallm-mcp

# MCP dependencies are included in core, just install with CUDA support
uv pip install -e '.[base,cuda124]'

# Verify installation
python -c 'from dnallm.mcp import server; print("MCP server dependencies installed!")'

Verification

Basic Verification

# Verify DNALLM import
python -c "import dnallm; print(f'DNALLM version: {dnallm.__version__}')"

# Verify core modules
python -c "
from dnallm import load_config, load_model_and_tokenizer
from dnallm.datahandling import DNADataset
from dnallm.finetune import DNATrainer
from dnallm.inference import DNAInference
print('All core modules imported successfully!')
"

Hardware Verification

# Verify PyTorch and CUDA
python -c "
import torch
print(f'PyTorch version: {torch.__version__}')
print(f'CUDA available: {torch.cuda.is_available()}')
if torch.cuda.is_available():
    print(f'CUDA version: {torch.version.cuda}')
    print(f'GPU: {torch.cuda.get_device_name(0)}')
    print(f'Memory: {torch.cuda.get_device_properties(0).total_memory / 1024**3:.2f} GB')
"

# Verify Mamba (if installed)
python -c "
try:
    from mambapy import Mamba
    print('Mamba: Available')
except ImportError:
    print('Mamba: Not installed')
"

# Verify Intel XPU (XPU build of PyTorch only)
python -c "
import torch
if hasattr(torch, 'xpu'):
    print(f'XPU available: {torch.xpu.is_available()}')
    if torch.xpu.is_available():
        print(f'GPU: {torch.xpu.get_device_name(0)}')
else:
    print('XPU: torch build has no XPU support')
"

Troubleshooting

CUDA Version Mismatch

Issue: Installed PyTorch CUDA version doesn't match system CUDA version

Solution:

# 1. Check system CUDA version
nvidia-smi
nvcc --version

# 2. Uninstall installed torch
uv pip uninstall torch torchvision torchaudio

# 3. Reinstall matching version
uv pip install -e '.[cuda121]'  # Choose based on actual situation

GPU Present but torch.cuda.is_available() Returns False

Issue: The GPU shows up in nvidia-smi, but PyTorch cannot see it.

Solution: Almost always a wrong PyTorch build. Check the wheel suffix:

python -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.is_available())"

  • 2.x.x+cpu — you have the CPU build. This happens with plain pip (pip ignores the per-CUDA indexes in pyproject.toml, and the default PyPI wheel for Windows is the CPU build) or when pip install -e '.[...]' re-resolved torch after a GPU torch was installed. Fix by reinstalling the GPU build LAST, e.g.:

pip install --force-reinstall --no-deps 'torch==2.11.0' \
    --index-url https://download.pytorch.org/whl/cu130
- 2.x.x+cuXXX but still False — driver too old for that CUDA build (cu130 needs NVIDIA driver >= 580); update the driver.

XPU: OSError: [WinError 126] Loading c10_xpu.dll / shm.dll

Issue: torch XPU build installed, but importing torch fails with "module not found"-style errors naming c10_xpu.dll or shm.dll.

Solution: The Intel SYCL/oneAPI runtime packages drifted out of sync with the torch wheel (common after a --no-deps force-reinstall of a different torch version). Either reinstall torch WITH its dependencies:

pip install --force-reinstall 'torch==2.11.0' \
    --index-url https://download.pytorch.org/whl/xpu

or install the exact runtime pins listed in the installed torch wheel's METADATA (<env>/Lib/site-packages/torch-2.11.0+xpu.dist-info/META-DATA), e.g. intel-sycl-rt==2025.3.2, intel-cmplr-lib-rt==2025.3.2, intel-openmp==2025.3.2, tbb==2022.3.1, intel-pti==0.16.0, onemkl-sycl-*==2025.3.1, ...

Mamba Installation Failure

Issue: mamba-ssm or causal_conv1d installation fails

Solution:

# 1. Install compilation dependencies
conda install -c conda-forge gxx clang ninja

# 2. Clear cache and reinstall
rm -rf .venv/lib/python*/site-packages/mamba_ssm*
rm -rf .venv/lib/python*/site-packages/causal_conv1d*
uv pip install -e '.[mamba]' --no-cache-dir --no-build-isolation

# 3. Or use installation script
sh scripts/install_mamba.sh

Dependency Conflicts

Issue: Dependency conflicts during installation

Solution:

# 1. Create new environment
conda create -n dnallm-new python=3.13 -y
conda activate dnallm-new

# 2. Use uv to resolve dependencies
uv pip install -e '.[base]' --resolution=lowest

Native Mamba Support

Native Mamba architecture runs significantly faster than transformer-compatible Mamba architecture, but native Mamba depends on Nvidia GPUs.

If you need native Mamba architecture support, after installing DNALLM dependencies, use the following command:

# For venv users: activate virtual environment
source .venv/bin/activate  # Linux/MacOS

# For conda users: activate conda environment
# conda activate dnallm

# Install Mamba support
uv pip install -e '.[mamba]' --no-cache-dir --no-build-isolation

# If encounter network or compile issue, using the special install script for mamba (optional)
sh scripts/install_mamba.sh  # select github proxy

Please ensure your machine can connect to GitHub, otherwise Mamba dependencies may fail to download.

Additional Model Dependencies

Specialized Model Dependencies

Some models use their own developed model architectures that haven't been integrated into HuggingFace's transformers library yet. Therefore, fine-tuning and inference for these models require pre-installing the corresponding model dependency libraries:

EVO2

EVO2 model fine-tuning and inference depends on its own software package or third-party Python library1/library2:

# evo2 requires python version >=3.11
# Install transformer torch engine
uv pip install "transformer-engine[pytorch]==2.3.0" --no-build-isolation --no-cache-dir
# Install evo2
uv pip install evo2
# (Optional) Install flash attention 2
uv pip install "flash_attn<=2.7.4.post1" --no-build-isolation --no-cache-dir
## Note that build transformer-engine and flash-attn package will cost much time.

# add cudnn path to environment
export LD_LIBRARY_PATH=[path_to_DNALLM]/.venv/lib64/python3.11/site-packages/nvidia/cudnn/lib:${LD_LIBRARY_PATH}

EVO-1

# Install evo-1 model
uv pip install evo-model
# (Optional) Install flash attention
uv pip install "flash_attn<=2.7.4.post1" --no-build-isolation --no-cache-dir

GPN

Project address: https://github.com/songlab-cal/gpn

uv pip install git+https://github.com/songlab-cal/gpn.git

megaDNA

Note that megaDNA weights stored at the Hugging Face can be accessed after requesting permission from the author.

Project address: https://github.com/lingxusb/megaDNA

git clone https://github.com/lingxusb/megaDNA
cd megaDNA
uv pip install .

LucaOne

Project address: https://github.com/LucaOne/LucaOneTasks

uv pip install lucagplm

Omni-DNA

Project address: https://huggingface.co/zehui127/Omni-DNA-20M

uv pip install ai2-olmo

Enformer

Project address: https://github.com/lucidrains/enformer-pytorch

uv pip install enformer-pytorch

Borzoi

Project address: https://github.com/johahi/borzoi-pytorch

uv pip install borzoi-pytorch

Some models require support from other dependencies. We will continue to add dependencies requirement for different models.

Flash Attention Support

Some models support Flash Attention acceleration. If you need to install this dependency, you can refer to the project GitHub for installation. Note that flash-attn versions are tied to different Python versions, PyTorch versions, and CUDA versions. Please first check if there are matching version installation packages in GitHub Releases, otherwise you may encounter HTTP Error 404: Not Found errors.

uv pip install flash-attn --no-build-isolation --no-cache-dir

Compilation Dependencies

If compilation is required during installation and compilation errors occur, please first install the dependencies that may be needed. We recommend using conda to install dependencies.

conda install -c conda-forge gxx clang

Verify Installation

Check if installation was successful:

# Test basic functionality
python -c "import dnallm; print('DNALLM installed successfully!')"

# Run comprehensive tests
sh tests/test_all.sh