Introducing System One Models (like Jev) into Your Data Platform
I tested Jev against an LLM on a practical data pipeline task: checking whether sources support extracted claims. The cost and speed looked promising. Here’s what the comparison showed, and what I’d test before automating review.
In this article
Close to two years ago, I joined the folks at Dagster and Hex for a webinar on introducing LLMs into a data platform. We were classifying narratives. The output was useful, but as the workload grew, so did the bill.
I've run into a related problem in Data Tech Market, a project where I'm mapping data tools and the debates around them. An LLM extracts claims from public sources. Before those claims enter the corpus, I need to read the source and check whether it actually supports what was extracted. That's a lot of review.
So when TypeSafe introduced Jev, I wanted to try it on that specific task. Could a model built for focused decisions help with a repetitive check inside a data pipeline? And what would I need to know before trusting it?
I compared Jev with GLM-5.3 using the same claims and supplied source text. The cost and response-time differences were promising. The harder question, whether the decisions were good enough to automate review, is still open.
Here's the full walkthrough, including the platform, a small request example, and the results.
The task: does the source support the claim?
Data Tech Market collects public material about data technologies. Part of the project is a catalog of companies and products. Another part is a corpus of attributed claims: what people say about those tools, and how those positions contribute to debates.
For example, semantic layers are at the center of plenty of discussion about agents and analytics. I want to make those arguments inspectable, with a source behind each claim.
The pipeline collects sources, extracts proposed claims, and sends them for human review. Accepted records then move into the corpus and downstream data marts.

Dagster orchestrates the work. The asset graph makes the collection, extraction, and proposal stages visible.

The distinction between a quote and a supported claim matters here. Finding the quoted words in a source is a useful check, but it doesn't establish that a summary preserves the speaker's meaning, scope, or qualifications.
In the walkthrough, I use a claim attributed to Ian Macomber about Ramp's investment in data modeling, dbt practices, semantic layers, and single sources of truth. The source passage connects that foundational work with what the company has been able to do with AI internally. The catalog summary adds more specific language about AI and text-to-SQL capabilities.

That's the kind of gap I want a check to catch. Does the supplied passage establish the whole claim, or only part of it? A narrow snippet may leave out relevant context. The answer has to be about the text supplied to the model, rather than everything the model might know about Ramp.
What Jev changes about the interaction
TypeSafe describes System One models as models that understand natural-language input and return typed decisions with probabilities. Jev is its first model in this category.
The basic interaction is straightforward: supply the context, define the possible answers, and ask a focused question. For this experiment, the output is one of three labels:
- Supported: the supplied text establishes the material claim, including its attribution, scope, and qualifiers.
- Contradicted: the supplied text directly conflicts with a material part of the claim.
- Not enough evidence: the supplied text neither establishes nor clearly contradicts it.
The useful part for a pipeline is that the answer arrives in a shape code can work with. I don't need an explanation paragraph when the next step needs a label.
TypeSafe also provides score and yes/no-style primitives. Those open up other possible tasks, such as classifying content or routing a record for review. But this experiment only tests source support.
And a valid label can still be a bad judgment. A predictable output format solves one part of the integration problem; evaluating the decisions remains our job.
A small example before the batch
I started with a simple Hex notebook using the TypeSafe Python SDK through OpenRouter.
The request packages the claim and a short source passage into the model's state. Then a Choice question supplies the instructions and the three possible answers.

In this example, Jev returned not_enough_evidence. The response also included a probability distribution and a separate confidence value. That distinction is worth keeping: those fields aren't interchangeable.
Given the narrow passage and the uncertainty in this response, I'd want the case reviewed. I haven't established a confidence threshold that would make automatic acceptance safe.
The OpenRouter log also shows how small an individual call can be. The pictured request reports a cost of $0.0000179 and provider latency of 123 milliseconds. That's one request, with one input size. It isn't the batch result or a promise about every future call.

Comparing the same task across two models
For the larger experiment, I used saved claims and their supplied source passages. Both Jev and GLM-5.3 received the same task and rubric. A Marimo app ran the requests and recorded their answers, confidence, request time, reported cost, and failures.
The original design started with 100 saved cases. The recorded run overview separates 31 additional development cases, three earlier pilot cases, and 66 held-out cases. Within the held-out group, the dashboard reports the main representative comparison on 64 cases and keeps two challenge cases separate.

This is an exploratory comparison. I didn't create new independent human reference labels for these cases. Historical review decisions provide context, but neither model is the reference judge.
That means I can compare their behavior, output coverage, speed, and cost. I can't turn their agreement into a measurement of accuracy or an estimate of how many wrong claims would pass an automated review rule.
What the comparison showed
For the 64 representative cases, the dashboard showed:
| Measure | Jev | GLM-5.3 |
|---|---|---|
| Valid answers | 64 of 64 | 57 of 64 |
| Supported | 57 | 52 |
| Not enough evidence | 7 | 5 |
| Contradicted | 0 | 0 |
| Median request time | 0.32 seconds | 2.30 seconds |
| Reported cost for 64 requests | $0.0075 | $0.2394 |
Among the 57 cases with valid answers from both models, 56 matched. That's 98.2% agreement, with one disagreement. Seven representative cases had no valid pair, so they aren't included in that agreement percentage.
The two separate challenge cases had valid pairs that agreed, too. Two cases are far too few to draw a broad conclusion about difficult inputs.
The speed and cost figures are encouraging for this workload. Jev's reported cost for those 64 requests was below one cent. GLM-5.3's was about 24 cents. Both amounts are small in isolation, but repeated semantic checks across a pipeline are exactly where those differences start to matter.
The timing is observed end-to-end API time through OpenRouter, not intrinsic model speed. The dashboard includes successful requests with invalid answers in the timing and cost comparison, and excludes cached replays from latency.
The output gap matters as well. Seven GLM-5.3 answers weren't usable in this representative comparison. I didn't investigate their cause in the video, so I wouldn't attribute all of them to one failure mode. In a working pipeline, though, unusable answers still need somewhere to go.
What I'd test before putting this into the review flow
I'm happy with this as a reason to keep exploring Jev. It returned usable answers across the representative cohort, largely agreed with the comparison model, and used less time and reported spend in this run.
But most labels were supported, and neither model returned contradicted in that cohort. That leaves an obvious question: what happens when the supplied text actually conflicts with the claim?
I'd stress-test that next, alongside claims with missing qualifications, ambiguous attribution, and passages that support only part of a summary. I'd also create independent reference judgments before measuring incorrect passes or deciding which cases could skip human review.
Confidence could become a routing signal, but I need to test it against this task. A high confidence number doesn't, by itself, tell me that a claim is safe to accept.
For now, the source-support check is an experiment on saved data. It hasn't replaced Data Tech Market's human review step.
Beyond this one check: semantic columns
The broader possibility is interesting. In “What Jev will do to data engineering,” Astronomer explores judgments stored as semantic columns, such as a lead's persona alongside the model's confidence.
That feels familiar to anyone who's stretched a CASE WHEN statement into a classifier. An LLM offers another way to interpret messy text, but its cost and response time can make running that interpretation on every row difficult.
Astronomer's article describes a possible flow where confidence helps route rows to the warehouse, a larger model, or a person. I like that direction. It gives the uncertain cases an explicit path instead of forcing every row through the same decision process.
It's also a separate use case, with its own evidence. Their job-title results don't validate my claim-checking task, and I haven't tested that cascade here.
What I take from this experiment is a concrete next step: find a repeated judgment in your pipeline, define exactly what the possible answers mean, and try it on representative inputs. Record failures and missing answers as carefully as successful ones.
Then decide what evidence you'd need before acting on the result. For my pipeline, the cost question looks promising. The trust question needs another experiment.