How many gene tokens does a single-cell foundation model need to reproduce its representation of a cell? And if several different inputs produce that representation, what can a selected gene set tell us about the model?

Experiments I've run show that a few hundred selected genes can approximate the full-input embedding, and that disjoint inputs can also work at sufficient-enough token budgets. I think this is an interesting interp question regarding model behavior and faithfulness, and also a useful basis for adaptive tokenization.

The main interpretive constraint is that reproducing an encoder is a narrower goal than preserving biological information. A compact panel can recover a representation and tell us about relationships between unique sets of genes.

Questions§

For a cell x, let G_x be its expressed genes within the model's vocabulary, and let S\subseteq G_x be a retained subset. Write z_S=f_\theta(T(x;S)) for the frozen encoder's output under native tokenization T, and z_{\mathrm{full}}=f_\theta(T(x;G_x)) for the reference (teacher) embedding.

For nonzero embeddings, fidelity is

\operatorname{Fid}_x(S)=\cos(z_{\mathrm{full}},z_S).

A set is \tau-sufficient when its fidelity reaches \tau. Zero-norm outputs need to be recorded separately because cosine is undefined for them.

Minimal sufficiency. What is the smallest sufficient input?

K_\tau^\star(x)=\min_{S\subseteq G_x} \big\{|S|:\operatorname{Fid}_x(S)\geq\tau\big\}.

This defines an optimum.

Substitution. Given a sufficient set S_A, can a different set reproduce the embedding? The strict test requires

S_A\cap S_B=\varnothing, \qquad |S_A|=|S_B|, \qquad \operatorname{Fid}_x(S_B)\geq\tau.

A larger disjoint replacement establishes substitutability at an additional token cost, rather than equal-budget substitution.

Degeneracy. How many distinct sufficient sets coexist? The paper asks this in terms of near-disjoint sets. For comparisons across experiments, I would also fix a gene budget and an explicit overlap criterion. A count without those constraints mixes representational redundancy with the amount of input each candidate receives.

Structure and downstream preservation. Does substitutability follow co-expression modules or the model's input gene embeddings? And do sufficient panels preserve the response to an intervention?

Methods§

The primary model is Tahoe-x1. The reference is each cell's complete expressed gene panel: a median of 1,645 genes and a maximum of 2,918, which are natively uniformly randomly selected (WITHOUT REPLACEMENT).

The draft uses greedy forward selection and Best-of-N search with N=128, testing fidelity thresholds \tau\in\{0.80,0.90,0.95\}. The encoder remains frozen. A separate search forbids genes in an existing sufficient panel and increases the budget available to a disjoint replacement.

The equal-size substitution experiment is different. At a budget of 64 genes, it compares replacements within modules, across modules, and at random. Modules are constructed either from co-expression or from input gene-embedding vectors (still fleshing this out further). In these swaps, the original gene's expression value is transferred to its replacement. That can create a gene–value pairing absent from the measured cell, so these results should be interpreted separately from measured-gene subset selection.

A reproducible implementation also needs to specify normalization, value encoding, special tokens (CLS), and whether expression values are recomputed after removing genes. Deleting tokens while retaining full-cell normalization and renormalizing a smaller panel are different interventions. Search comparisons should report encoder-call counts and runtime as well as the number of genes retained.

Also examined Stack, whose embeddings depend on surrounding cells. I exclude it from the quantitative comparisons here because that context needs to be controlled before comparing selection budgets across models.

Results§

Smaller inputs reproduce the full embedding§

At a fidelity threshold of 0.95, greedy selection finds sufficient sets with a median of 288 genes; Best-of-128 finds a median of 512. Both are substantially smaller than the median full panel of 1,645 genes.

Median cosine fidelity versus gene budget in Tahoe-x1. Greedy and Best-of-128 cross 0.90 around 256 genes; random subsets lag. Greedy is shown through 512 genes, as in the paper.
Figure 1. Median fidelity to the full-input embedding, digitized from the paper's Figure 1. Lines connect the plotted points on a logarithmic budget axis. The greedy curve stops at 512 genes in the source.

Greedy finds smaller panels at each reported threshold. This is a comparison of retained tokens, not total computation. A search that spends more model evaluations to find a shorter input can still cost more overall.

The distinction between an optimum and a search result matters here. Every sufficient set found gives an upper bound on K_\tau^\star(x). Failure to find a smaller one does not establish minimality. Even a panel from which no single gene can be deleted may be larger than a different sufficient combination.

Disjoint panels also work§

Table 2 reports that median fidelity for the best disjoint sets reaches 0.90 at a budget of 256 genes and 0.95 at 512 genes. It also reports zero irreducible cases at 0.95 among the 47 cells over the searched budgets.

Median fidelity of disjoint Tahoe-x1 gene panels rises with budget. It crosses 0.90 at 256 genes and 0.95 at 512 genes, approaching 1 at 1,024 genes.
Figure 2. The Tahoe-x1 median curve from the paper's Figure 2, digitized from the PDF. The thresholds agree with Table 2: 256 genes for median fidelity of at least 0.90 and 512 for at least 0.95. These are not per-cell minimum budgets or evidence of an equal-size replacement for each cell's smallest found panel.

A gene may occur in one compact explanation of an embedding while another explanation excludes it entirely. So I guess its importance depends on which other genes are available to the selector.

Co-expression improves substitution scores§

Within-module substitutions based on co-expression score higher than cross-module or random substitutions. The corresponding gap is much smaller when modules are constructed from the model's input gene embeddings.

At 64 genes, co-expression grouping gives within, cross, and random substitution fidelities of 0.270, 0.053, and 0.050. Input-embedding grouping gives 0.077, 0.062, and 0.053. All are below the lowest sufficiency threshold of 0.80.
Figure 3. Reported substitution scores from Table 2 at a budget of 64 genes. The table does not specify a mean or median for these values. The dashed line marks the lowest sufficiency threshold used in the study.

The relative difference seems informative. The absolute scores limit the conclusion: 0.270 remains far below 0.80, so this experiment does not show that within-module swaps generally preserve the embedding. It shows that this grouping produces better substitutions under this protocol.

Nor does weak performance from input-embedding clusters settle what gene relationships the encoder knows. Static lookup vectors and contextual representations are different objects. The result compares two grouping procedures, not all possible ways of extracting biological structure from the model. Trying to make this better.

What these results establish§

The evidence supports compressibility and degeneracy under a particular encoder and similarity measure. It does not establish that the retained genes are biologically essential, that omitted genes are uninformative, or that the selected inputs preserve every downstream task.

Cosine similarity ignores vector magnitude. It could miss changes that matter to a downstream model using raw embeddings. Aggregate fidelity can also hide failures in rare cell states.

For perturbation work, the target is often a difference between conditions:

\Delta z=z_{\mathrm{treated}}-z_{\mathrm{control}}.

Close agreement with each endpoint does not automatically preserve a small difference between them. In single-cell experiments, this comparison also needs appropriate control and treated populations; the notation does not imply repeated destructive measurements of the same cell.

I would therefore treat easy compression as a property to investigate, not an explanation by itself. Correlated genes may provide redundant measurements of a transcriptional program. An encoder insensitive to important distinctions may also be easy to compress. The fidelity test alone does not separate those cases.

What this means§

A practical tokenizer probably needs a cheap selector. Expensive discrete search could provide training targets, but running a full encoder pass before every selection would undermine the intended inference saving. Evaluation should include selection, preprocessing, and encoding costs.

There is also an information-budget distinction. A computational selector may inspect the full measured expression vector before choosing tokens. This can reduce encoder work without reducing the measurements required to obtain the cell. A smaller experimental assay needs a panel chosen without access to unmeasured genes in each new sample. Probably have to trace K/V activations and attention compute, FLOPs or something.

The result worth pursuing is lower total cost while preserving the biological comparisons the model is meant to support + a better understanding of how these models treat gene relationships. Reproducing the full embedding gives a useful first target. Perturbation responses and rare-state discrimination would provide stronger tests of whether the tokenizer retained the information that matters.