← Back to blog

An L2 Risk Classifier That Reads Logits Instead of Generating

Risk assessment in Open Cradle has two layers. L1 is deterministic rules, done in under 10 ms. L2 is an LLM classifier. The final level is finalLevel = max(L1, L2): whatever the model says, a yellow from the rules never gets downgraded.

The problem was L2. The model wrote JSON under a grammar: category, risk, confidence, reasoning. On Qwen3 4B on M-series hardware that took 6–7 seconds per ticket, and the confidence field was whatever the model felt like writing. There was a number in the response; there was no meaning in it.

Hypothesis

Classification is a closed question. For a closed question the model does not need to write anything: one forward pass and the next-token distribution over the allowed answers is enough. Reading logits is a known technique, and I invented nothing here. What interested me was what would break when I moved it into a product running on small local models.

The decide() primitive

Each answer option gets a single-token alias: letters A, B, C… for choice and boolean questions, digits for a scale. The runner wraps the prompt in the model's chat template, leaves the assistant turn open and returns the probability mass on each candidate.

Three quantities per answer:

  • probabilities — the mass, renormalised over the options;
  • coverage — the sum of the mass on the options before renormalisation;
  • confidence = p(chosen) × coverage.

The ticket text comes first in the prompt. Its evaluated prefix is reused across the questions of one request, so the second and later questions are cheap.

import { decide } from '@cradle/core/decision'

const res = await decide(runner, {
  // The ticket goes first: its evaluated prefix is shared by every question.
  state: ticketText,
  // System turn: definitions of risks and categories. No "return JSON" lines.
  context: logitsContext(operatorPrompt),
  questions: {
    risk: {
      type: 'choice',
      instructions: 'What is the risk level of this message?',
      choices: {
        green: 'safe to answer automatically',
        yellow: 'an operator must review the answer',
        red: 'always needs approval'
      }
    },
    commitment: {
      type: 'boolean',
      instructions: 'Does the message ask us to take on a financial or legal commitment?'
    },
    urgency: {
      type: 'score',
      instructions: 'How urgent is this message?',
      min: 0,
      max: 5
    }
  }
}, { debug: true })

res.method                          // 'logits' | 'generated'
res.answers.risk.probabilities      // { green, yellow, red } — sums to 1
res.answers.risk.coverage           // mass on A/B/C before renormalisation
res.answers.risk.confidence         // p(chosen) × coverage
res.answers.commitment.probability  // P(yes)
res.answers.urgency.value           // expected value over 0..5
res.answers.urgency.mode            // most probable grade
res.calibrated                      // false — raw model output

If the runner cannot expose logits (an external API), decide() falls back to the generated path: an answer under an enum schema, probabilities: null.

Timing

VariantTime
JSON generation, Qwen3 4B, M-series6–7 s
Logits, 5 questions987 ms
Logits, 6 questions (with red flags)1674 ms
+ the "who handles it" question on the shared prefix+410 ms

What broke

1. The prompt asked for JSON. The operator prompt for the old path ended with "return JSON". On the logits path the model dutifully opened its answer with {, and coverage sat at 0.00–0.45. The lines about output format were removed from the context; the definitions of risks and categories stayed.

2. Reasoning models start with <think>. If that is the most probable next token, the runner appends an empty thinking block and reads the distribution after it. On Qwen3 this fires on every question.

3. Coverage 1.0 does not mean correct. A default context of "answer with a single letter" pushed coverage to 1.0, but the 0.6B and 1.7B models started picking the last option almost every time. Rolled back. Coverage measures format compliance, not correctness.

4. One three-way question missed red. On "please sign the contract extension", Qwen3 4B answered yellow with p = 1.00. The fix is three separate boolean red flags: a financial or legal commitment; deletion, access rights, production; third parties' personal data. Any "yes" with P ≥ 0.5 makes the verdict red. A model that misjudges the level usually still recognises the fact when asked directly.

5. Risk is not picked by argmax. A false green (an unreviewed auto-reply) costs far more than a spare yellow (an operator takes a look). Hence thresholds:

function pickRisk(p: Record<'green' | 'yellow' | 'red', number>) {
  if (p.red >= 0.3) return 'red'      // red as soon as P(red) reaches 0.3
  if (p.green >= 0.8) return 'green'  // green only when P(green) reaches 0.8
  return 'yellow'                      // everything in between
}

6. The fallback cannot return green. If coverage on risk or category drops below 0.5, the classifier falls back to the old generative path. But a verdict from the fallback can never be green: a prompt injection on Phi-4-mini produced exactly that false green. A model refusing a closed question is itself a sign that the message is unusual.

Measurements

44 hand-labelled tickets, Q4_K_M quantisation. Columns: risk accuracy, false greens, missed reds (yellow instead of red, so the ticket still reaches an operator), over-escalations, category accuracy.

ModelRiskFalse greenMissed redOver-escalatedCategory
Qwen3 4B95%00298%
Gemma 3 4B84%04393%
Qwen2.5 3B84%00784%
Phi-4-mini80%10891%
Llama 3.2 3B59%09948%
Qwen3 1.7B59%010868%
Qwen3 0.6B41%002648%

The caveat is mandatory: the labels are mine, the set is small. The numbers are good for comparing models against each other, not as a measure of absolute quality. The single false green in the whole table is that prompt injection on Phi-4-mini, which is where rule number six came from.

Limitations

  • The probabilities are not calibrated. A confident model is wrong at p = 1.0, and no threshold catches that. The response says so: calibrated: false.
  • Positional bias in small models is not compensated at all.
  • The red flags are tuned for one model. On another, both the wording and the threshold have to be re-checked on the set.

Next: temperature scaling per question, with labels taken from operator decisions; the ECE metric; averaging over permutations of the options against positional bias. Only after that can the thresholds be chosen honestly rather than by eye.

This article was created in hybrid human + AI format. I set the direction and theses, AI helped with the text, I edited and verified. Responsibility for the content is mine.

← Back to blog