A dinosaur needs an answer
The CLM authors' demo plays a faithful Python clone of Chrome's dinosaur game, stepped at 60 frames per second. Each decision is one request: a line of text describing the screen, a question, and three options.
Nothing here asks for prose. The program needs one of three words, and it needs it before the cactus arrives. That is the shape of most AI calls inside software: pick one, from a menu someone wrote in advance.
Request format from the CLM repository's examples/t_rex. Decision latency, from request to the game loop collecting the answer: a median 16.5 ms for CLM on a local RTX 4090 and 149.8 ms for Jev through its hosted API. The model calls alone took 2.6 ms and 131.9 ms. All numbers from the authors' results files.
Three ways to choose
The writer is a generative LLM. It reads the situation and the menu, then writes the answer one token at a time, and the program hopes the string names an option.
The judge reads the situation together with each option and scores every pair. Jev's architecture is unpublished; this joint reading is how Sebastian Raschka builds a Jev-like system, so it stands in for Jev here. It is accurate, and it re-reads every option for every new situation.
The lookup is CLM. It reads the situation alone and compares it with option vectors computed before the game began. Watch where the reading happens in each lane.
Now make the menu a thousand long
A web agent choosing which link to click may face hundreds of options per page. The CLM blog measured all three approaches with a 200-token state, running CLM and constrained decoding on one RTX 4090 and calling Jev's hosted API.
From 1 candidate to 1,024, CLM went from 36 ms to 44 ms. Jev went from 131 ms to 579 ms. Constrained decoding with the same Qwen3-8B, the writer forced to spell a valid option, reached 4.3 seconds.
The rest of this story explains why one bar barely moves.
Heights are on a log scale. Self-reported, 5 trials each; Jev measured through its hosted API.
The output layer you already know
Every language model already ends with a choice. Its final hidden state h is compared, by dot product, against one stored row per vocabulary token: 151,936 rows for Qwen3. A softmax over those scores is the next-token distribution.
So a writer already chooses by looking up a vector in a table. Its trouble is the table: fixed at training time, made of word fragments, so naming a team can take several sequential passes. Qwen3's tokenizer splits “Refunds” into Ref and unds: two passes, each waiting on the last.
The trouble with the denominator
A softmax divides each score by the sum over every row. When the rows are a vocabulary, every document on the web, or every image ever captioned, that sum is too large to compute at each training step.
The fix, found again and again, is to compare the right answer against a small handful of wrong ones and let those stand in for the rest. Train the right row to beat a few random rows, many times over, and it learns to beat them all.
Thirteen years of the same trick
In October 2013 two papers appeared within weeks of each other. Microsoft's DSSM trained a query tower and a document tower to pick the clicked page out of five, four of them random. word2vec's negative sampling trained each word to tell its true neighbour apart from a handful of random ones, one yes-or-no decision per pair.
The 2018 CPC paper gave the loss its name, InfoNCE. Sentence-BERT made two-tower search practical in 2019. CLIP in 2021 used every other caption in a batch of 32,768 as its negatives. CLM in 2026 does the same with an action menu.
Each pedestal shows what the denominator was replaced with.
A tower that writes the classifier
The CLIP paper describes its text encoder as “a hypernetwork which generates the weights of a linear classifier.” Each label's description becomes one row of the output layer.
That sentence is the key to CLM. The action head writes rows into the table from chapter one. A new action needs no retraining, only a description, embedded once and stored. The writer's vocabulary is frozen; this table is yours.
The wording of a row matters. CLIP gained 1.3% on ImageNet just by writing “A photo of a {label}.” instead of the bare label. Writing good action descriptions is CLM's prompt engineering.
Meaning as a place
An embedding is a list of numbers, which is also a position. Scale every vector to length one and they all live on the surface of a sphere, where the dot product is simply the cosine of the angle between two points.
Here are our six support teams, embedded once by the action head. Close means compatible. Choosing means finding the nearest team.
A toy in three dimensions so it can be drawn; CLM-8B's space has 512. Two-encoder models can also leave the two sides in separate regions of the sphere: Liang et al. found this “modality gap” in CLIP and traced it to the two encoders rather than the modalities. Only the ranking matters.
A score is a sum of agreements
Our running ticket: “My invoice was charged twice and nobody answers the phone!” Each column is one feature. Multiply the ticket's height by the team's height in each column, add the products, and you have the cosine.
Billing agrees on billing and refund. Escalation agrees on urgency and anger. Billing wins because its agreement is larger.
Our toy backbone has 8 named features so you can see why. Real features have no names.
A giant reader, two small heads
CLM-8B is a frozen Qwen3-8B run as an embedding model. Its vector is read at the last token, the only position in a causal model that has seen the whole input: 4,096 numbers.
Two heads map those numbers into a shared 512-dimensional space. Each is a small MLP, 4096 → 1536 → 1536 → 512. Drawn to scale by parameter count, the reader is about 870 times the size of one head. Only the heads are trained.
From the blog's layer sizes: 9.4M weights per head, 18.9M for both, which matches the 75 MB checkpoint at 32-bit precision. The blog's “20M-parameter projection head” counts both heads together.
Why the heads are needed
A frozen LLM's vectors already encode meaning, but not this relation. Asked “Who wrote the play Romeo and Juliet?”, raw Qwen3-8B embeddings rank William Shakespeare third among the candidate answers.
After CLM's first training stage, the same backbone with two trained heads ranks him first, at 54.2%. The heads do not learn what Shakespeare is. They learn to put a question next to its answer.
Example from the CLM blog. It does not list the other candidates, so they are drawn unlabelled.
One pass, one lookup
A decision is now three moves. The ticket and the question go through the frozen reader once. The state head turns that into one point. That point is compared with every cached action row in a single matrix-vector product, and a softmax turns the scores into probabilities.
The action rows were computed before the ticket arrived.
A batch is a grid
Training takes B pairs that belong together: a state and the action actually taken in it. Score every state against every action and you get a B × B grid with the right answers on the diagonal.
Each row is a B-way choice and so is each column; the loss, bidirectional InfoNCE, averages both. Every off-diagonal cell is a free negative. Press train and watch real gradients raise the diagonal.
Temperature decides who gets blamed
For one row, the gradient pushes each wrong option down in proportion to its softmax probability. The towers show that share for the ticket's five wrong teams.
At a high temperature the push is spread across all five. At a low one it lands almost entirely on the most confusable options: here Refunds and Escalation, which sit at nearly the same cosine from this ticket, split it about evenly. CLIP learns its temperature, starting at 0.07 and capped at 0.01. CLM starts from the same value, and its released checkpoint sits at that cap.
Wang & Liu (2021) call the trade-off the uniformity–tolerance dilemma.
Look-alikes, then the real job
Random negatives teach topic: a refund request is not a Mario jump. Hard negatives teach distinctions: a double charge is not a payment-method question. In our toy, adding them raises the smallest gap between a state's true action and its look-alike from about zero to between 0.24 and 0.47.
CLM-8B trains in three stages: about 60M question–answer pairs, 30M synthetic hard negatives, then about 1M agent trajectories, each step a state–action pair, mixed with 40% replay of the first stage. Hard negatives used too early peak at 62.4% and overfit; after pre-training they reach 69.2%.
Platform areas are proportional to example counts. Ablation numbers are self-reported.
Pay for the menu once
Because the two sides never meet before the last step, the action side can be computed in advance and stored. Over T decisions against K options, a judge reads T·(S + K·A) tokens. CLM reads K·A once, then T·S.
Drag the sliders. The judge's tower grows with the menu; CLM's grows only with the number of decisions.
What the cache cannot skip
On one RTX 4090, going from 3 to 50 cached actions added at most 0.3 ms. A state the server has seen before came back in 0.6 ms. A brand-new state cost about 28 ms every time, because the 8B reader still has to read it.
The saving needs a menu that is bounded, stable and written in advance. Generate fresh candidates on every call and there is nothing to reuse.
clm-serve timings from the CLM README, server-side p50, cache off → on. Self-reported.
A probability among these
Billing holds 0.83 of the probability. We are about to add a fourth option, “Invoices and payments”, which means the same thing. What happens to Billing?
One vector, decided in advance
The state head never sees the options, so it must squeeze the situation into one vector that will serve every option it might meet. A judge re-reads the situation for each option and can look for exactly the detail that option needs.
This is the price of being cacheable, and a plausible reason CLM trails Jev on plain sentiment classification: 82.90% against 96.47% on IMDb, in an independent test by Sebastian Raschka.
Read the receipts
Both models survived all five dinosaur runs with the same score, 697. That tie is guaranteed: with no deaths on a deterministic clone, every 60-second run covers the same distance. The fine print matters more. A physics planner writes “Safe… Best.” or “Unsafe… Collision.” into every option's text, and a safety layer checks every answer before the dinosaur moves.
CLM chose the planner's best move 65.8% of the time; Jev, 98.7%. Across five runs the safety layer stepped in 4,883 times for CLM, 364 of them overriding a pick labelled unsafe and 4,519 correcting an answer that no longer fit the frame it landed on, and 28 times for Jev. The system survived; the safety layer did much of the work.
From the CLM repository's T-Rex results files. CLM ran locally; Jev ran over the network, so the speed gap is real but not like-for-like.
Use both: retrieve, then judge
Search settled this years ago. The Sentence-BERT paper measured the gap: finding the most similar pair among 10,000 sentences takes about 65 hours with a cross-encoder and about 5 seconds with two towers. The library's documentation recommends the combination: retrieve with the towers, then re-rank the shortlist with the cross-encoder.
The same cascade fits decisions. Let CLM cut a thousand options to ten in a few milliseconds, then let a judge read those ten carefully.
Where the two sides meet
Every design in this story is a choice about one thing: how early the situation and the options are allowed to read each other.
The writer mixes everything and can say anything. The judge mixes each pair. Late interaction, as in ColBERT, keeps a vector per token and compares them at the end. The lookup meets only in a single dot product, which is why it can be cached and why it misses fine detail.
A decision becomes a lookup
Jev is named after William Stanley Jevons, who noticed in 1865 that as steam engines learned to do more work per ton of coal, Britain burned more coal, not less. TypeSafe's bet is the same for decisions: every order-of-magnitude drop in cost unlocks far more uses.
CLM pushes that further by turning a decision into one pass and one lookup, at the cost of calibration and fine detail. Below, train one yourself.