MicroGlot

A taxonomy-informed sparse DNA foundation model for microbial genomics.

MicroGlot is a 23-layer decoder-only mixture-of-experts transformer pretrained on 378.3 billion nucleotides from 3.70 million sequences across 99,700 microbial species, spanning bacteria, archaea, fungi, protists, viruses and plasmids. It encodes the taxonomic hierarchy as hyperbolic (Poincaré) embeddings and uses them both as an input token and to steer expert routing.

Paper: A Taxonomy-Informed Sparse DNA Foundation Model for Microbial Genomics
Code: github.com/athanzli/MicroGlot

Models

Model Input Use it when Load with
MicroGlot DNA and its species most of your sequences have a known species from_pretrained("athanzli/MicroGlot", ...)
MicroGlot-plain DNA most of your sequences have no known species from_pretrained("athanzli/MicroGlot", subfolder="plain", ...)

A species is known if it is one of the 99,700 pretraining species (check with tokenizer.has_species(name)). For taxonomic prediction tasks, use MicroGlot-plain, as conditioning on the species would leak the label.

Model details

Architecture decoder-only transformer, next-token prediction
Layers / hidden size 23 / 1024
Parameters 2.98 B, of which 479 M are active per token
Mixture of experts 312 routed experts across layers (U-shaped), top-1 routing plus a shared expert
Species conditioning MicroGlot uses a 32-d Poincaré embedding as an input token and to modulate expert routing, and MicroGlot-plain uses none
Context 8,192 tokens (about 43 kb)
Tokenizer byte-pair encoding, vocabulary 8,192

Installation

MicroGlot is tested on Linux with an NVIDIA GPU and requires FlashAttention-2 (flash-attn), whose rotary position-embedding kernel it was trained with.

conda create -n microglot python=3.12 -y
conda activate microglot
pip install torch==2.8.0 "transformers>=4.51.3,<5.19"
pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.8.3.post1/flash_attn-2.8.3.post1%2Bcu12torch2.8cxx11abiTRUE-cp312-cp312-linux_x86_64.whl

Tested with Python 3.10 to 3.13, PyTorch 2.7 to 2.13, transformers 4.51.3 to 5.18 and flash-attn 2.7.4 to 2.8.3.post1. For another Python or PyTorch version, install the matching flash-attn wheel from the flash-attn releases.

Usage

MicroGlot

import torch
from transformers import AutoModel, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("athanzli/MicroGlot", trust_remote_code=True)
model = AutoModel.from_pretrained(
    "athanzli/MicroGlot", trust_remote_code=True, dtype=torch.bfloat16
).to("cuda")

sequences = ["ATGAGTAAAGGAGAAGAACTTTTCACTGGAGTTGTCCC", "TTGACAGCTAGCTCAGTCCTAGGTATAATGCTAGC"]
species = ["Escherichia coli", "Bacillus subtilis"]

inputs = tokenizer(sequences, species=species, padding=True, return_tensors="pt").to("cuda")
with torch.no_grad():
    outputs = model(**inputs, output_hidden_states=True)

hidden_states = outputs.hidden_states   # 24 x [2, length, 1024]; [0] is the token embeddings, before any decoder layer
  • species= takes one name per sequence, or one name for all of them. Names are NCBI Taxonomy scientific names (September 2025), e.g. Clostridioides difficile. Use species-level NCBI names; tokenizer.species_names lists all 99,700. Case and extra spaces do not matter, and _ or - count as spaces. An unknown name raises a KeyError listing similarly spelled names; check that a suggestion is the same organism.
  • outputs.hidden_states[k] is the output of decoder layer k (1 to 22); [0] holds the token embeddings, and [23] is last_hidden_state, the output of layer 23 after the final normalization. All outputs line up with input_ids; padded positions have attention_mask 0.

Sequences without a known species

If a small portion of your sequences have no known species, consider discarding them, so that every remaining sequence is given its exact taxonomy embedding. To keep them instead, use the Species-encoder to infer their species embeddings from the DNA and fill these gaps. Continuing the MicroGlot example, pass None as their species.

model.load_species_encoder()   # downloads the Species-encoder (6 GB) and attaches it to MicroGlot

inputs = tokenizer(sequences, species=["Escherichia coli", None], padding=True, return_tensors="pt").to("cuda")
with torch.no_grad():
    outputs = model(**inputs, output_hidden_states=True)   # the second sequence's species is inferred

If most of your sequences have no known species, using the Species-encoder to infer their species is not recommended due to performance considerations. We recommend using MicroGlot-plain instead.

Long sequences

The context is 8,192 tokens (about 43 kb). To encode a longer sequence, one viable way is "chunk and encode", by cutting the sequence into windows that fit the context and encoding each window. Continuing the MicroGlot example, the code below fills each window to the full context of 8,192 tokens.

genome = "ATGAGTAAAGGAGAAGAACTTTTCACTGGAGTTGTCCC" * 3000   # stand-in for a 114 kb sequence
# the sequence's DNA tokens, without [BOS] and [EOS] (verbose=False skips the length warning)
ids = tokenizer(genome, add_special_tokens=False, verbose=False)["input_ids"]
# windows of 8,190 tokens, decoded back to DNA; the tokenizer then adds [BOS] and [EOS], 8,192 in all
windows = [tokenizer.decode(ids[i:i + 8190]) for i in range(0, len(ids), 8190)]

window_embeddings = []
with torch.no_grad():
    for window in windows:
        inputs = tokenizer(window, species="Escherichia coli", return_tensors="pt").to("cuda")
        hidden_states = model(**inputs, output_hidden_states=True).hidden_states   # 24 x [1, length, 1024]
        window_embeddings.append(torch.cat(hidden_states).mean(dim=1))             # [24, 1024], mean over tokens
sequence_embedding = torch.stack(window_embeddings).mean(dim=0)                    # [24, 1024], mean over windows

A single embedding of the whole sequence is then the average of the window embeddings.

MicroGlot-plain

import torch
from transformers import AutoModel, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("athanzli/MicroGlot", subfolder="plain", trust_remote_code=True)
model = AutoModel.from_pretrained(
    "athanzli/MicroGlot", subfolder="plain", trust_remote_code=True, dtype=torch.bfloat16
).to("cuda")

sequences = ["ATGAGTAAAGGAGAAGAACTTTTCACTGGAGTTGTCCC", "TTGACAGCTAGCTCAGTCCTAGGTATAATGCTAGC"]

inputs = tokenizer(sequences, padding=True, return_tensors="pt").to("cuda")
with torch.no_grad():
    outputs = model(**inputs, output_hidden_states=True)

hidden_states = outputs.hidden_states   # 24 x [2, length, 1024]; [0] is the token embeddings, before any decoder layer

MicroGlot-plain takes no species; everything else works as for MicroGlot.

Limitations

  • MicroGlot was trained on inputs of up to 8,192 tokens.
  • Only the 99,700 pretraining species can be given by name. Other species count as unknown (see Models).

Citation

Li, A. Z., Wang, S., Cheng, S., Du, Y. & Liu, R. A Taxonomy-Informed Sparse DNA Foundation Model for Microbial Genomics. bioRxiv (2026). https://doi.org/10.64898/2026.09.22.753215

@article{li2026microglot,
  author  = {Li, Athan Z. and Wang, Shiyuan and Cheng, Shupeng and Du, Yuxuan and Liu, Ruishan},
  title   = {A Taxonomy-Informed Sparse {DNA} Foundation Model for Microbial Genomics},
  journal = {bioRxiv},
  year    = {2026},
  doi     = {10.64898/2026.09.22.753215},
  url     = {https://www.biorxiv.org/content/10.64898/2026.09.22.753215v2}
}
Downloads last month
550
Safetensors
Model size
3B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support