Key takeaways
- Decision models take a piece of state and typed questions, and return a bounded choice, an ordinal score or a yes/no with a probability distribution, without generating text. The task is zero-shot classification; the packaging is what is new.
- Hosted and open-weight versions both exist, so a decision model can run inside the SMO or on the O-Cloud. That removes the remote round trip and the data residency objection, and opens a research path for some slower Near-RT decisions, subject to measured loop latency.
- A probability over labels is not the probability that an action will succeed. Gating a network change needs a quantitative risk model, hard constraints and the cost of being wrong alongside the model's judgement.
- Calibration is not supplied by the interface or by the training method. It has to be measured on operator data with a clean evaluation design, and monitored afterwards.
A new interface for an old task
Most of the attention in AI over the last few years has gone to models that write. You give them a prompt and they produce text, and if what you actually wanted was a decision, you then parse that text, validate it and hope the model picked one of the options you offered rather than something it made up.
In September 2026 a group of models arrived that skip the text. They take a state, which can be a text string or a structured object, together with a set of typed questions, where each question is a Choice between options you define, a Score on ordered levels you define, or a yes/no. The model returns, for each question, a probability over the allowed answers. Because the output schema is closed, the model cannot return an answer type that does not exist. It can still return an allowed answer that is confidently wrong.
On 15 September 2026 TypeSafe announced Jev as its first public model and opened early access through a hosted API; no weights were released. On 17 September the Kev project published its first release, a LoRA adapter and pointer head on Qwen2.5-0.5B, with Apache 2.0 covering those two components and the base model, which is downloaded separately, under the Qwen licence. On 18 September Convai Innovations published the initial Laya checkpoint on Hugging Face under Apache 2.0, and added its multilingual and typed-decisions checkpoints to the same repository the following day. Other projects followed, including wrappers that read option probabilities from the logits of existing open models. These are publication dates, and they say nothing about when any of the underlying research began.
What these models share is an interface, not an architecture. Jev's documentation describes scoring all questions in one query, but its internals are not public. Laya documents batched inference. At least one open wrapper runs one forward pass per question. Compatible APIs expose bounded Choice, ordinal Score and yes/no questions, while the exact output fields and inference strategies vary. A valid schema makes a result easier to consume in code; it does not establish that the answer is correct or that its uncertainty is calibrated.
Nor is the task itself new. Classification with probabilistic output is decades old, and zero-shot classification with label sets supplied at inference time was already documented in 2019. TypeSafe's own CEO agreed with a commenter that Jev is essentially a zero-shot classifier. What is new is the packaging: any typed question answered without the caller training a classifier, many questions scored against one state, and a probability returned in a form that code can threshold directly. Zero-shot means no task-specific training on the caller's side. It does not mean reliable answers to arbitrary questions, which is why the old lessons about classifiers still apply.
What the numbers in the answer mean
Every automated decision in a RAN sits somewhere on a spectrum that runs from acting now, through collecting more information first, to doing nothing at all. A model that simply returns "sleep the cell" gives you no way to place the decision on that spectrum, whereas a model that returns a distribution over "sleep", "keep" and "not enough information" at least tells you how divided its own judgement is.
This is close to the argument we make about our own models, where a coverage estimate without a confidence interval hides whether it should be trusted. But two kinds of uncertainty are in play, and keeping them apart matters more than anything else here. A model assigning a high probability to the label "sleep" is reporting how strongly the state matches that option. It is not estimating the probability that a carrier shutdown will preserve coverage and slice performance. Those are different prediction targets, and a peaked distribution over labels says nothing about the second one. Prediction uncertainty, forecast intervals and the causal effect of an action are three separate things, and a control loop needs all three.
The vocabulary differs between implementations. In Jev's documentation, Choice and Score return a probability distribution plus a separate confidence statistic, while the yes/no question returns only its probability. Confidence is described as a function of the shape of the distribution, with no single universal formula. Score is the expectation over ordinal level indices, not simply a selected level and not a continuous physical measurement, which is one more reason to keep arithmetic and KPI comparisons in code.
Where decision models could fit
None of the uses described here has been demonstrated in a RIC. The releases and benchmarks available so far are general purpose, and each use would need its own evaluation on operator data before it is trusted.
The Non-RT loop is the plausible starting point. The Non-RT RIC runs control loops of one second and longer, most rApp loops run on minutes, and a decision model answering in tens or hundreds of milliseconds fits that budget, whether hosted or deployed as a container next to the rApp. Within that loop, four uses are worth testing, and all four share one property: the rApp computes the quantities and hands the model a bounded judgement.
The first is gating proposed actions. Suppose an energy saving rApp proposes a carrier shutdown. Before the change is applied over O1, a decision model scores the proposal against a compact summary the rApp has prepared: the load trend as a named bucket, the categories of recent alarms, whether a planned event overlaps the window. The model's answer is one evaluated signal. It sits alongside a quantitative network risk model, hard operating constraints and the cost of a wrong action, and the application's policy combines them into automatic approval, deferral or human review. The model cannot replace those checks; its contribution is a cheap semantic read of context the numeric model does not see.
The second is classifying text that reaches an rApp from outside the RAN. Most of what an rApp consumes over R1 is structured: alarms carry a probable cause and a perceived severity, counters have names and values. But some context arrives as free text, such as the additional-information field of an alarm, a planned-work notice from the operations team or a maintenance window description, and an rApp may need a bounded judgement about it, for example whether a notice concerns a cell in the energy saving cluster or whether a maintenance window overlaps a planned action. Routing alarms to teams and prioritising tickets remain functions of the operator's fault management and ITSM processes, not of the RIC. For the long tail of such free-text questions, the ones nobody has labelled data for, a text-first decision model is a natural tool.
The third is checking a language model. In intent-driven operation an operator writes "reduce energy in cluster X without affecting the URLLC slice" and a language model turns it into policies. A decision model judging whether the generated policy matches the intent is a useful additional signal, but it is itself fallible, so it belongs next to deterministic policy and constraint checks rather than in place of them.
The fourth, and weakest, is a coarse value-of-measurement decision. Before running an expensive collection campaign, a Score question over a summary of the current model uncertainty and the campaign's cost could act as a pre-filter. The real decision still needs the numeric model, but the example shows the shape.
Local deployment adds a fifth possibility. The Near-RT RIC runs between 10 ms and 1 s. The Laya model card reports 32.8 to 39.5 ms for a single question on a Tesla T4 and 193 to 464 ms on CPU; its 7.2 ms per question figure is amortised across a ten-question batch that takes 72.3 ms in total, so it is not a single-decision response time. Local inference removes the remote round trip and creates a plausible research path for some slower xApp decisions, for example whether a pattern of E2 indications matches a known failure signature. Suitability depends on the complete loop deadline, including feature preparation, queueing, inference, E2 communication and action execution, and it has to be established by measuring tail latency under representative load. A fast average inference time is not enough.

Where they do not fit
The current general-purpose releases have not demonstrated suitability for per-slot DU control below 10 ms, and nothing about them suggests they will. Models built for the real-time loop have different timing properties and are a different subject.
The more important limit is numbers. TypeSafe's documentation for Jev 1.13, reviewed on 17 September 2026, lists counting, numeric representations, date arithmetic, irrelevant context, adversarial content and consistency between question forms as known failure modes, and anticipates fixes in later versions. These findings are specific to that version, but the engineering consequence holds for every current release: semantic decision models should not replace validated models for energy attribution, coverage prediction or throughput forecasting. Arithmetic, dates, KPI comparisons and physical constraints belong in code or in specialised models, and the decision model should receive only the information its bounded judgement needs.
Context follows the same rule. Accuracy falls as the state fills with material unrelated to the question, and context windows are bounded, so raw performance tables and traces do not belong in the state. The application filters and summarises first.
Untrusted text is a concern the class shares with language models. The state is treated as data rather than as something hostile, so content written to steer the model, including injected instructions, can move the answer. Alarm text, tickets and release notes can contain material the operator did not write, which means a decision built on them needs the same input controls a language model would need.
Consistency across question forms cannot be assumed either. The documented behaviour includes a yes/no question and its negation not summing to one, and a threshold tuned on one question type not transferring to another, so each decision should be asked one way with any identities enforced in code.
Finally, the benchmarks that exist are task specific and say nothing about the RAN. A reproducible benchmark run on 29 September 2026 scored Jev at 79.2 percent on 3,080 Banking77 test messages; with a threshold fixed on validation data, it accepted 51.2 percent of test items with 3.68 percent error among accepted predictions. On the open side, Laya's base English checkpoint scored 0.362 on the typed-decisions benchmark against a 0.461 majority-class baseline, while its checkpoint fine-tuned on that benchmark's training split scored 0.766. Both results show the same pattern: performance depends on the task, the checkpoint and the acceptance policy, and none of it transfers to alarms or change requests without being measured there.
Treat it as a governed component
Deployment is now a choice rather than a constraint. A hosted model brings a vendor's evaluation effort and no infrastructure to run, at the cost of operational data leaving the operator and a version that moves unless it is pinned. An open-weight model can run on the O-Cloud or inside the SMO, keeps the data in place, can be pinned by file hash, and can be fine-tuned or recalibrated on operator data. Open weights alone do not establish compatibility with a given platform's packaging and services, nor procurement acceptance, but they make the operator the owner of the calibration rather than a consumer of someone else's claim.
Where labelled history already exists, compare against a small supervised classifier trained on it. Whether the supervised model wins depends on the task, the labels, the data distribution and the deployment constraints, which is exactly why the comparison should be run rather than assumed. The zero-shot model earns its place on questions nobody has labels for and on questions that change faster than a training cycle, and its own deferral rate is a reasonable signal for which questions have become worth labelling.
In either case, treat the decision model as a governed component of the rApp or xApp: record its version, evaluate it on operator data before it is trusted, and monitor its behaviour in production. Integrate with the deployment's model management services where they apply. The fact that no training happened on the operator's side does not change any of this.
Calibration deserves a clean design, because probability output and calibration-oriented training do not guarantee calibration on new data, and post-hoc methods such as temperature scaling help only when they are fitted and tested properly. The Kev release notes give a concrete example at small scale: on 1,350 held-out in-distribution questions the checkpoint reports 0.799 accuracy and an expected calibration error of 0.065, falling to 0.031 after temperature scaling with a single fitted parameter. In-distribution is the operative word, and the same procedure has to be repeated on the operator's own questions. Define the correct outcome for each question. Fit any calibration layer and select action thresholds on development and validation data, then freeze them before evaluating an independent test set. Check reliability by probability band and by operating regime, together with the error among accepted decisions and the fraction of cases accepted. If the model proves overconfident, recalibrate and reselect the policy according to measured risk; moving the acceptance threshold changes how many decisions are automated, not how honest the probabilities are. Evaluate deferred cases as well as accepted ones wherever labels can be obtained, since outcomes observed only after selected actions bias the assessment. Then keep monitoring after deployment, because the network changes under the model.

What we make of it
The interesting thing about decision models is neither speed nor price, and it is not any particular vendor. It is that they make a probability over a bounded answer set the primary output rather than something extracted from text afterwards, and for the judgement step of a control loop that is a sensible default. It is close in spirit to the position we take in our own research, where an estimate of a network quantity, such as attributed energy or predicted coverage, carries its uncertainty as part of the answer rather than as an afterthought. The difference is that those models estimate quantities, while decision models estimate labels.
The uses worth testing today are triage and gating in the Non-RT loop, with the quantities computed in the rApp, the decision model supplying one evaluated signal among several, and the model deployed where the data already lives. The Near-RT case is a research question with a measurable answer.
The division of labour is the part we are confident about: specialised models compute the network quantities, a bounded semantic classifier supplies judgements, and application code applies the policy.
