By John Siciliano, Knut Melvær
Published
We made a ludicrously fast docs agent with Jev, Knowledge Bases, and gpt-oss-120b by moving four decisions off the LLM. Here's the (not-so) secret sauce.
Jump to section
We have all gotten used to watching agents “think.” You ask it a question, watch a spinner spin, and some time after you’ll see the answer unfold in the UI. We put up with this because the answers are presumably good and generative AI-backed agents still feel somewhat new.
We used to put up with this for websites for the same reason. Images would load one pixel row at a time on our 56k modems. Now that high-speed Internet is common, we don’t want to wait for a website to load anymore.
Agents are headed the same way, and that’s what's exciting about Jev, the new “System One model” that came out recently and had everyone hyped up (including myself). It lets you use a model trained with “world knowledge” to make decisions really fast and way cheaper than a Large Language Model (LLM).
So for the past week, I've been exploring ways to use Jev to make agents faster.
We created a blank project called “wicked fast agent” (we guess it was a foregone conclusion?) and the results we kept having beat every expectation we had.
The agent responds blazingly fast and is generally accurate.
Here are my findings and methods to optimize agents to the point they compete with a webpage’s load time.
Jev made LLMs the wrong tool for certain jobs
Jev is a “decision”/“classification” model. Like LLMs, it is trained on a lot of data in order for it to have knowledge about the world, but unlike LLMs, it does not produce a bunch of text. It’s optimized to receive structured documents that describe a state and a set of typed questions. In practice, it looks a bit like this:
{
"model": "jev-latest",
"state": "How do I set up live preview with Next.js?",
"questions": {
"entry": {
"type": "choice",
"instructions": "Which knowledge base entry most likely contains the answer?",
"criteria": {
"frameworks/nextjs": "next-sanity setup, App Router, fetching and caching",
"visual_editing/live_preview": "enableLiveMode(), query stores, Live Content API",
"studio/configuration": "defineConfig, workspaces, environment variables",
"none_of_these": "No entry covers the question."
}
}
}
}{
"model": "jev-1.13.0",
"answers": {
"questions": {
"type": "choice",
"choice": "visual_editing/live_preview",
"confidence": 0.43,
"probabilities": {
"frameworks/nextjs": 0.42,
"none_of_these": 0,
"visual_editing/live_preview": 0.58,
"studio/configuration": 0
},
"stats": {}
}
},
"usage": {
"input_tokens": 412,
"output_tokens": 66
},
"request_id": "playground_1085515399b0ffe495c92b9f6ab3a4c3158",
"evaluation_time_ms": 71.75401900167344
}And while you can have an LLM make a decision like this, Jev is the proper tool for the job. Similarly, an 18-wheeler truck can deliver a single pizza, but it would be a lot faster and cheaper to load it into a sedan.
In my preliminary testing we started to see these results:
- Faster: 4.7x to 13.5x faster in my testing
- Cheaper: 44x to 346x cheaper in my testing
- Can't go off-script: It can only put probability scores on the options you give it, never invent new ones. It can still pick the wrong option, which is what the confidence score is for.
- Know when it’s not confident: Every decision lets you know how confident it was in that choice
So if you use it as a model router in your agent and ask it to decide between using Sonnet and Haiku to handle a given request based on its complexity, and you can trust it to pick one of those two 100% of the time. Compared to asking an LLM such as Sonnet 5 the same question, it's 6.0x faster, 116x cheaper, and it was just as accurate in my tests.
Let that sink in.
We now have two general use cases:
- Start using Jev to make decisions we previously relied on LLMs for.
- Start using Jev to make decisions we would have never handed to an LLM because they would be impractically slow and expensive.
These two use cases are the foundation of the four methods that make my agent ludicrously fast.
Making a docs agent fast
The demo is a documentation agent. You ask it a question about Sanity and it answers from our documentation. We have “precompiled” our documentation into Sanity Knowledge Bases, which extracts key info from all the articles, organizes it into entries and creates an index optimized for agentic lookup. Of course, you don’t have to use our product for doing this, but it is a convenient and fast way to set up a content backend for agents.
With the obligatory product pitch out the way, let’s look at what the agent’s job in this demo is.
Answering a question takes three steps:
- Decide. Which entry has the answer, which model should write it, and whether to look into the knowledge base, query the articles directly, or neither.
- Fetch. Pull the entry from the knowledge base, or the article from the query API.
- Answer. The chosen model reads the retrieved content into its context and streams a reply.
In a typical agent, we would hand the first step to an LLM with tool calls, one after another, each one waiting for the former to be resolved. Yes, this is a waterfall, similarly to what we’re used to when building frontends that rely on APIs to provide content for the render.
In my agent, Jev makes the Step 1 decisions instead, while you are typing.
Four interesting Jev findings
With these four methods, our agents would pass the Google Core Web Vitals if they applied to agentic experiences: the median answer finishes in 0.5 seconds, well under the 2.5-second bar for a "good" Largest Contentful Paint.
Here is how the times compare in my testing:
Decision | LLM | Jev | Saved |
|---|---|---|---|
Entry choice | 1,328ms (Haiku) | 282ms | 1,046ms |
Speculation | n/a | 0ms at enter (done while typing) | 1,828ms |
Model routing | 818ms (Haiku) | ~300ms | 511ms |
MCP choice | 829ms (Haiku) | 170ms | 659ms |
1. Finding the right context (saves 1,046ms)
The first optimization is using Jev to decide which entry in our knowledge base matches the user's prompt. We provide context to the agent with Sanity Knowledge Bases, which extracts key info from our content, organizes it into a knowledge base, and creates an index. The index is how the agent knows where to look for content.
Using a knowledge base is in itself an optimization, since we’re “pre-compiling” facts for agents so that they don’t need to do broad searches and interpret and build the context at inference time. It’s slightly similar to the “statically built” vs “server-rendered” site comparison from the web-perf world.
studio/configuration
defineConfig, workspaces, environment variables, auth settings, project setup, and deployment.
studio/customization
Custom form components, component overrides, tools, structure builder, document actions, and theming.
visual_editing/live_preview
enableLiveMode(), query stores, Live Content API integration for preview, and real-time content streaming.
visual_editing/overlays
enableVisualEditing(), click-to-edit overlays, stega encoding, data-sanity attributes, and content source maps.And while this index improved the speed for LLMs, it took Haiku 1,328ms just to decide which entry contains the necessary content. The agent still needs to fetch that entry and find the answer within. And of course with an LLM, especially small models, the risk of the model inventing an entry, i.e., hallucinating, is ever present
The solution is clearly having Jev decide which entry is relevant.
My evals showed this happens 4.7x faster and 44x cheaper than Haiku, with great accuracy.
entry: choice("Which knowledge base entry most likely contains the answer to `question`?", {
"studio/configuration": "defineConfig, workspaces, environment variables…",
"studio/customization": "Custom form components, component overrides, tools…",
"visual_editing/overlays": "enableVisualEditing(), click-to-edit overlays, stega…",
// …one option per entry in the index
none_of_these: "No entry covers the question.",
})Because Jev returns probabilities, it knows when a question spans two entries. For "How do I set up live preview with Next.js?" it gave Next.js 0.74 and Live preview 0.19, so the agent reads both. An LLM tool call just returns a list of paths, with no confidence to decide when a second entry is worth reading.
Let's take this a step further.
2. Deciding before the user hits enter (saves 1,828ms)
Jev is named after William Stanley Jevons, of Jevons paradox:
An economic phenomenon said to occur when technological improvements that increase the efficiency of a resource's use lead to a rise, rather than a fall, in total consumption of that resource.
In other words, since Jev is relatively cheap and fast, we can just throw more stuff at it. Like instead of waiting until the user hits enter to decide which entry to select from the knowledge base, we have Jev speculate which entry is relevant as the user types, so by the time they hit enter, we already know what entry is relevant, can fetch it and have it ready for when the agents need to answer.
This is an example of a use case we would never rely on an LLM for.
Speculative decisions save us 1,328ms compared to using an LLM on enter, and 282ms compared to using Jev on enter. While inexpensive (about $0.00008 per speculation, so you could speculate 44 times for the price of one Haiku decision), we still recommend applying debouncing to conserve resources, such as waiting for 250ms of no activity before speculating.
As a part of this step, we prefetch the entry, which saves us around 500ms. Now when the user presses enter, the agent already has the context it needs to answer the question.
The LLM that reads the entry and responds to the prompt will depend on your domain. Regardless of the domain, the goal is to choose the smallest model while still maintaining accuracy and instruction following.
For more complex domains, such as ones with code examples, a relatively larger model is needed, but for simpler domains, you can get by with a small model, like one with just 32B parameters.
For my demo agent, I’m using openai/gpt-oss-120b on Cerebras, as it needs to reason about a more complex corpus. In my tests it was also the most accurate of the fast models (20 of 22 answers fully grounded in the docs) and the fastest, with the first token arriving in about 350ms.
But we don’t have to be tied to just one model. Let’s dynamically route prompts using Jev.
3. Dynamic model routing (saves 511ms)
A small, 32B parameter model can tackle questions like "Can I use external sources with Sanity Knowledge Bases?", but will struggle with "Write me a GROQ query that finds semantic similarities to basketball".
We can dynamically choose a model which is a structured decision, making it a perfect use case for Jev.
{
"questions": {
"model": {
"type": "choice",
"instructions": "Which model should answer this question?",
"criteria": {
"small_model": "Factual lookups answered directly from the docs",
"large_model": "Writing or debugging code, multi-step reasoning"
}
}
}
}Provide two choices and some instructions on when to choose each and in around 300ms, Jev will select the appropriate model, ensuring users get quality results and you aren't footing an Opus bill for almost regex-viable answers.
One caveat: in my demo, routing made answers cheaper, not faster. gpt-oss-120b on Cerebras is already the fastest model we tested, so sending prose questions to a smaller model on a slower host added about half a second at the median (roughly 1.0s instead of 0.6s). It still cut the cost of those answers by about 4.5x. For the 0.5-second demo we turned routing off. In production, where cost matters more than the last few hundred milliseconds, I'd keep it on.
There's one last decision our agent needs to make.
4. Decide which retrieval method to use (saves 659ms)
By now you are probably seeing the meta pattern. Is there a decision with a constrained set of options to be made? Use Jev.
A documentation agent largely relies on the retrieval-augmented generation (RAG) pattern to answer technical questions precisely. You can combine different search and look up strategies depending on the complexity and nature of the question. Sometimes you need very specific context, and sometimes you need broader.
We typically load our agents up with a multitude of MCP servers, and the agent decides which one(s) it needs for the given task. In my testing, we gave Jev a choice between two endpoints (or both, or neither) from Sanity Context:
- Knowledge base lookups, for questions about documentation, plans, and policies.
- Structured content queries, for filtering content such as products.
Jev makes this choice in 170ms compared to Haiku's 829ms and Sonnet's 2,299ms.
Better yet, these decisions don't have to happen one after another. Jev answers several questions in a single request: entry choice, model routing, and MCP selection together took about 340ms, only about 60ms more than entry choice alone.
{
"model": {
"choice": "large_model",
"probabilities": {
"large_model": 0.99,
"small_model": 0.01
}
},
"source": {
"choice": "knowledge_base",
"probabilities": {
"knowledge_base": 0.85,
"neither": 0.08,
"dataset": 0.04,
"both": 0.03
}
}
} Find the decisions hiding in your prompts
We can anticipate the pressure to make our agents faster/better/cheaper. With models like Jev, we can get ahead of this by starting to think about what agentic tasks can be turned into decisions, rather than an open-ended prompt. For example:
- Open-ended: "Figure out how to handle this technical question"
- Decision-based:
- “Which knowledge base entry has the answer?”
- “Does this task need the large model?”
- “Knowledge bases, structured content queries, both, or neither?”
It works outside agents too. "Review this pull request" becomes "Does this change a public API?" and "Does the description match what the code actually does?"
Put together, my agent gives full answers in 0.5 seconds at the median, and three out of four in under 0.9 seconds. It costs about a third
Granted, these numbers come from a week of testing on one corpus, so treat them as indicative and not a true benchmark.
And here’s hoping for agents that answer as fast as the page they’re on.
More from Engineering