Concepts
Composite Scoring
Split one fuzzy judgment into atomic Jev questions, then combine the numbers in code with weights you control.
Quick answer
Ask Jev several narrow questions about the same state. Combine the returned probabilities in your code with weights you can diff. Official guidance keeps that sum in application code, and points at a classical model only when you have labels for the combined outcome.
A "good ticket" is several facts: the customer asked for something you can do, the message is not a legal threat, and nobody is asking for a refund above a limit you already computed. Jev can answer each fact. Your code can add them up.
One request, several questions
Questions in a single System One call are answered against the same state and do not see each other's answers. That is the right batch for a composite. You pay for the state once. See Patterns for the official parallel-questions cookbook, which measures that batching on a long article.
answers = response.answers
quality = (
0.5 * answers["request_is_actionable"].noul
+ 0.3 * (1 - answers["needs_legal_review"].noul)
+ 0.2 * (1 - answers["refund_demanded"].noul)
)
if quality >= 0.8 and answers["request_is_actionable"].noul >= 0.7:
auto_handle(ticket)
else:
review(ticket)The weights are the policy. Changing 0.5 to 0.2 is a code review, not a prompt rewrite. The second condition stops a high average from hiding one atomic failure. Confidence-gated routing is that second check.
When a Score is enough
Use a Score when the product really is an ordered scale with 2–10 levels you can describe, and you do not need to explain the mix later. Use a composite when support, risk, or ranking will ask "which part failed?"
When you have labels
If you later know which tickets were actually good, the Nouls are features. Official docs point at a classical model on top of those features, and the autoresearch cookbook is their worked example of proposing questions and fitting a regressor. That cookbook is theirs to run. The split stays the same: Jev emits numbers, a model you train on labels combines them, and the combination is still not a prompt.
Lead scoring on this hub is a single Score. A composite is the next step when one rubric is doing too many jobs.
FAQ
Why not ask Jev for the final score directly?
You can, with a Score question. A composite keeps each reason visible, so a product change is a weight change. A single Score hides the mix inside one distribution.
Should the weights live in the prompt?
No. Official guidance puts the weights in code. The model answers the atomic questions. Your program decides how they add up.
Sources
- Composite scoringTypeSafe · accessed 2026-09-30 · documentation
- How to build with TypeSafeTypeSafe · accessed 2026-09-30 · documentation
- Autoresearch feature discoveryTypeSafe · accessed 2026-09-30 · documentation