Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

typed-lm

New Deterministic inference · Adapter training · Apache-2.0

Structured decisions in a single forward pass.

typed-lm turns dense decoder models — Llama, Qwen2, Qwen3, Mistral, Gemma, Gemma2 and Gemma3 — into a typed semantic-routing API. Send a state and typed questions; receive booleans, choices and scores your code can branch on. No text generation, no parsing.

7
dense model families
3
question primitives
4
training methods
1
forward pass per request

What does typed-lm do?

typed-lm is a Rust workspace with a Jev-compatible HTTP server and a trainer. The server loads a dense decoder checkpoint once and answers typed questions from the logits at a single decision position, instead of generating text.

---
accTitle: One request, one forward pass
accDescr: A client sends a state and questions; the server evaluates them in one batched forward pass and returns typed answers.
---
flowchart LR
  client["Client"]:::neutral
  request["state + questions"]:::primary

  subgraph model["typed-lm-serve"]
    direction TB
    prefill["shared prefill"]:::accent
    batch["batched decision positions"]:::accent
  end

  answers["typed answers<br/>noul · choice · score"]:::success
  code["your code<br/>branch · sort · route"]:::success

  client --> request --> prefill --> batch --> answers --> code

  classDef primary fill:#ede9fe,stroke:#7c3aed,color:#3b0764,stroke-width:1.5px
  classDef accent fill:#dbeafe,stroke:#2563eb,color:#0c4a6e,stroke-width:1.5px
  classDef success fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:1.5px
  classDef neutral fill:#f4f4f5,stroke:#a1a1aa,color:#18181b,stroke-width:1.5px

Why typed decisions?

Text-generation APIs force you to coerce a generative model into emitting structured output and then parse it back. typed-lm removes that mismatch: the model is scored with a restricted cross-entropy at the decision position, and the API returns a typed value plus a calibrated distribution.

⚡
One forward pass

All questions in a request share a prefill and are evaluated in one batched pass. Adding questions barely changes latency.

🎯
Calibrated by training

LoRA, QLoRA and full training optimize the exact decision-position loss the server reads at inference.

🧩
Jev-compatible

Drop-in compatible with the Jev contract: noul, choice and score, combinable in one call.

📦
Servable artifacts

FP8/FP4 quantization and full/from-scratch checkpoints are served directly by the same binary.

The API you will call

curl -s http://127.0.0.1:8080/v1/systemone \
  -H 'Content-Type: application/json' \
  -d @examples/request_mixed.json
{
  "model": "typed-lm",
  "answers": {
    "refund_eligible": { "type": "noul", "noul": 0.87 },
    "responsible_department": {
      "type": "choice",
      "choice": "logistics",
      "probabilities": { "billing": 0.05, "logistics": 0.9, "product_support": 0.05 },
      "confidence": 0.85
    },
    "urgency": {
      "type": "score",
      "score": 1.2,
      "legend": { "0": "Routine", "1": "Urgent", "2": "Emergency" },
      "probabilities": { "0": 0.2, "1": 0.4, "2": 0.4 },
      "confidence": 0.2
    }
  },
  "usage": { "input_tokens": 512, "output_tokens": 4 }
}
Next step: follow the Quick start to build the server, send your first request and train a LoRA adapter.

Who is it for?

  • Platform teams that need fast, auditable decisions instead of generated text.
  • ML engineers in Rust who want deterministic inference and a training and quantization pipeline.
  • Jev users who want a self-hosted, open-source implementation of the same contract.

The Jev contract,
on your own model and your own hardware.

Build with us

typed-lm is open source (Apache-2.0) and advances crate by crate. Contributions are welcome — from datasets and prompts to CUDA backends.