When you need software to decide something about a piece of text — which support queue a ticket belongs to, whether a message is spam, how positive a review is — the industry reflex over the last few years has been to reach for a large language model. It works, but it is often slow, expensive, and awkward to parse. In September 2026 two new models reframed the problem: Jev, a commercial model from TypeSafe, and Laya, an open-source alternative from ConvAI Innovations. Both belong to an emerging class of non-generative decision models. This article compares them with traditional LLMs so you can pick the right tool for the job.
What Are 'System One' Models?
The name borrows from Daniel Kahneman's idea of 'System 1' thinking — fast, intuitive judgments — as opposed to slow, deliberate reasoning. A System One model does not generate sentences. Instead, you give it a context, a question, and a fixed set of allowed answers, and it returns a typed result with a probability for each option. Jev exposes three question shapes: a choice (pick one of up to 255 options), a score (a position on a defined scale), and a yes/no probability. Because the output is structured, your code can consume it directly with no prompt-parsing, no JSON repair, and no risk of the model inventing a label that does not exist.
LLMs: The Flexible Generalist
Large language models remain the most capable option when the task is open-ended. They can classify text zero-shot from a plain-English prompt, explain their reasoning, handle labels that change on every request, chain multiple steps, and work across text, images, and audio. The trade-offs are latency, cost, and control: a classification call can take hundreds of milliseconds to several seconds, pricing scales with tokens, probabilities are poorly calibrated, and the free-text output must be parsed and validated. For high-volume, high-frequency classification where you only need a label — not an essay — an LLM is often doing far more work than the problem requires. That is precisely the gap Jev and Laya aim to fill.
Jev: The Commercial Benchmark
Released by TypeSafe on 15 September 2026 as the first 'System One' model, Jev is a closed, hosted API. It is priced at roughly $0.042 per million input tokens with output unmetered, and responds in about 70 to 500 milliseconds per request (independent tests report a median closer to 236 to 276 ms). Its headline strength is zero-shot performance on hard, high-cardinality tasks: it scores around 0.87 on Banking77's 77-way intent classification and 0.910 on AG News, and its probabilities are relatively well calibrated (calibration error around 0.144). TypeSafe trains it with a technique it calls Reinforcement Learning for Calibrated Decisions (RLCD) and claims up to 193.6x speed and 444.6x cost advantages over language models, while noting those figures sit at the high end of real-world gains. The catch: it is text-only, closed-source, and you are dependent on a vendor and a per-call bill.
Laya: The Open-Source Challenger
Three days later, on 18 September 2026, ConvAI Innovations released Laya on Hugging Face under a permissive Apache-2.0 licence, with the source code on GitHub. It ships as a small, self-hostable model built on bidirectional encoders — ModernBERT-large (about 421M parameters) for English and mmBERT-base (about 322M parameters) for over 100 languages. It is roughly an 808 MB download, runs in under 1 GB of memory, and returns an answer in about 33 milliseconds on a modest Tesla T4 GPU with no per-call cost. On low-cardinality tasks it is excellent, actually beating Jev on AG News (0.950) and DAIR Emotion (0.595). The decisive caveat is that Laya is best treated as a fast base you fine-tune on your own data rather than a drop-in zero-shot replacement.
"The right model is not the most powerful one — it is the smallest one that reliably solves your actual problem."
— Ascylla R&D
Head to Head: Accuracy, Speed and Cost
On raw speed, Laya leads at around 33 ms versus Jev's 236 to 276 ms, and it is free to self-host against Jev's $0.042 per million tokens. On low-option tasks the two are close, with Laya often ahead. The gap opens on high-cardinality, zero-shot problems: on the 77-label Banking77 benchmark Jev scores about 0.870 while Laya's out-of-the-box checkpoint drops to roughly 0.425 — below the task's majority-class baseline. Laya's reported 0.766 on a typed-decisions benchmark came from a checkpoint fine-tuned on that same dataset, so it reflects specialisation rather than general zero-shot skill. In short: Jev wins zero-shot flexibility and many-option accuracy; Laya wins latency, cost, privacy, and control once you have labelled data to tune it.
The Calibration Caveat Nobody Should Skip
Both models are fast, but fast confidence is not the same as correct confidence. Independent adversarial testing of Laya — scenarios with negation, retractions, and buried evidence — found zero-shot accuracy around 40%, lower than simply returning a constant answer, and over half of its high-confidence (0.90+) predictions were wrong, with answers shifting roughly one in five times when choice order was shuffled. Jev is better calibrated but not immune. The lesson for any production deployment is the same: validate on your own representative data, recalibrate probabilities, and never trust a confidence score you have not measured. This is exactly the kind of evaluation discipline that separates a demo from a dependable system.
Which Should You Choose?
Reach for an LLM when questions or labels change constantly, when you need explanations or multi-step reasoning, or when inputs include images and audio. Choose Jev when you need strong zero-shot accuracy on many-option decisions immediately, with no training data and no infrastructure to run — and you are comfortable with a hosted, metered API. Choose Laya when you have labelled examples and GPU capacity, and you value millisecond latency, zero per-call cost, data privacy, and full ownership of the weights. And do not overlook the oldest answer: if your labels are stable and you have enough examples, a conventional fine-tuned classifier or even a few rules can still be the simplest, cheapest, and most reliable option.
How Ascylla Can Help
Choosing between an LLM, a hosted decision model like Jev, and a self-hosted model like Laya is ultimately an engineering and economics decision — one that depends on your volume, latency budget, accuracy bar, data sensitivity, and team. At Ascylla we design, benchmark, and ship these systems end to end: building honest evaluation harnesses on your real data, fine-tuning open models like Laya on your taxonomy, integrating hosted APIs where they earn their cost, and wiring the result into production with monitoring and calibration built in. The goal is never the trendiest model — it is the one that reliably moves your numbers.
The arrival of Jev and Laya signals a healthy maturing of applied AI: not every text decision needs a billion-parameter generalist. As this category grows, the advantage will go to teams that match the model to the task and measure the results rigorously. Whether that means an LLM, a System One model, or a classic classifier, the winning move is the same — choose deliberately, verify honestly, and keep your options open.
Official Resources & Links
Jev & System One — TypeSafe AI announcement
System One — TypeSafe AI documentation
Laya — model card on Hugging Face
Laya — source code on GitHub
Large Language Models — overview (Wikipedia)

