When an AI agent needs to decide whether to query an order-status service, search delivery-policy docs, or ask a customer for an ID, it does not need a paragraph of prose. It needs a score. A new approach to serving decision-focused language models, detailed by researchers at SGLang, separates prompt design from runtime execution to make those scores faster and cheaper to produce at scale.
The core insight: a causal language model already holds a distribution over possible next tokens at the position where an answer would begin. If the model can express a judgment at that boundary, application code can read the relevant scores directly. No generation required.
Pointwise vs. setwise: what the model sees
The researchers distinguish between two prompt formulations. Pointwise scoring evaluates each candidate action independently, producing one Yes/No row per option while keeping other candidates hidden. Setwise scoring exposes all candidates together and returns a single selection row, such as A/B/C. Both can use explicit label scoring, but the scores carry different meanings.
"The distinction is what each judgment can see, not which API serves it," the authors write. "Choose the formulation based on the task and the model's training, including whether candidates should influence one another."
Open-Jev's public request compiler provides a concrete pointwise example. It builds independent candidate prompts from a shared context-and-question prefix, one proposed answer, and an instruction to answer Yes or No. That creates a serving workload with repeated state, multiple independent evaluations, and very small outputs.
Why a dedicated scoring interface matters
One-token generation with logprobs works as a baseline, but its top-k response can omit a label the application needs. SGLang's /v1/score endpoint lets callers declare required labels explicitly through label_token_ids, alongside the query and items. The runtime can then skip token sampling, avoid unnecessary logprobs for input tokens, gather label scores in batches, and reduce GPU-to-CPU transfers.
This label-selective extraction does not remove the vocabulary projection or full-distribution normalization. Both scoring and modern one-token generation can return a result directly from prefill, so generation does not inherently require an extra forward pass.
Multi-item scoring reuses the shared query
Pointwise candidates repeat the same state and question. Single-item scoring (SIS) processes them as independent logical sequences. Multi-item scoring (MIS) explicitly reuses the shared query within a request and restricts each candidate's attention to that query and its own tokens.
"Continuous batching and prefix caching can help, but batching alone does not guarantee shared computation," the authors said. MIS provides that guarantee. In plotted workloads, latency grew much less with candidate count and offered load when MIS was enabled.
The researchers recommend measuring achieved throughput and tail latency at the intended load, with enough client capacity to avoid mistaking a load-generator limit for a server limit. Benefits vary by model and configuration.
Getting started with the Score API
A pointwise example returns one Yes/No row for each possible action. For SIS execution, launching a standard server is straightforward. For MIS on supported models, the server requires FlashInfer, the radix cache disabled, chunked prefill set to -1, and the --enable-mis flag. Label token IDs must be derived from the checkpoint's tokenizer rather than copied between models.
For prompts that expose all candidates together, a separate Fused-Choice/Setwise study showed comparable latency when both endpoints share the same server configuration. Which path is faster depends on the model and offered load.
Why this matters for IT and research professionals
Decision workloads that need structured signals rather than generated text are common in agentic systems, recommendation pipelines, and routing logic. Explicit score-only serving removes the fragility of hoping a label appears in top-k logprobs. MIS offers a measurable latency reduction when candidate counts grow, but the gains are not uniform across architectures. The practical takeaway is to benchmark your own model, prompt formulation, and load profile rather than assuming one serving path fits all use cases.
Your membership also unlocks: