Rendered at 04:06:29 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
clintonb 1 days ago [-]
> Ultimately, we believe that incorporating Jev-driven diagnosis into an SRE agent’s workflow is a significant step toward effectively combining System One and System Two models.
Why? What’s the end goal? Lower cost? Faster analysis?
I configured an agent to respond to pages via Slack. It has access to ClickStack (for telemetry), Kubernetes, and GitHub. It runs one of the Sonnet models. The cost is so low, responding to less than a dozen pages a day (mostly from a very sensitive error count alarm) that swapping got Jev makes zero sense. This is especially true if the results are less trustworthy.
victor9000 20 hours ago [-]
using jev is a significant step towards using jev
s_Hogg 18 hours ago [-]
No-one ever got fired for using Jev
cyanydeez 17 hours ago [-]
Jev: The Decider
yimingsu 13 hours ago [-]
Hello! One of the blog's co-authors here. Faster analysis and lower cost is the goal. We think that a Jev-driven pipeline can be a quick fast-path filter for a lot of alerts that are recurring or even false-positives.
We can also add a confidence output to this pipeline, so we invoke a heavier reasoning agent (like Sonnet) when something suspicious deserves a closer look.
beebmam 1 days ago [-]
It seems to me that Jev is designed for when many quick decisions, with low input context, need to be made. SRE work is probably best suited for very few critical decisions that need to be made, with high input context.
yimingsu 12 hours ago [-]
That's a good point. At the same time, a lot of an SRE's daily work is also processing recurring alerts that are of less significance. We believe that a Jev-driven pipeline can be a helpful to SREs in quickly analyzing them.
codeduck 10 hours ago [-]
This is what an operations engineer or first line support does.
An SRE goes and fixes the cause of the recurring alerts.
amne 21 hours ago [-]
I switched to a decision model for my local "AI" voice control thingy. faster-whisper produces the text then onnx scores all my HA entities to determine which one I'm talking about .. then it lists all its actions and a second scoring round to determine what action I want. some regex to extract numbers if I'm saying things like "water the lawn for 10 minutes". It's pretty dumb when compared to what an LLM can do but considering you now have semantic scoring "at your finger tips" without resorting to levenshteins or other crude algorithms like that it's insane how smart it can look.
And to answer the obvious question: because latency. This thing turns on the light or water valve in under 100ms on a laptop from 2001 with 8gb ram .. all running locally on the laptop.
jazzyjackson 3 hours ago [-]
Consider looking at ChatScript for these kind of formulaic language matching -> variable extraction. It’s deterministic and instantaneous compared to any LLM
yimingsu 13 hours ago [-]
This is amazing! We are also fascinated by how fast Jev is. It's very much a breath of fresh air in the age when LLMs tend to take a much longer time and reason about tasks.
nikolovv 15 hours ago [-]
under 100ms on a 2001 laptop answers the why-not-llm question. what happens when the top two entities score close together on a short command like lights off, do you just take the first or ask back?
yehosef 18 hours ago [-]
repo?
mrkn1 17 hours ago [-]
[dead]
tinthedev 19 hours ago [-]
Using Jev for SRE seems like such a backwards prospect. You don't need split-second velocity, you don't need costs savings. What you DO need and Jev provides not at all - is auditability. You want to know WHY and HOW the calls were made, as with humans so with AI. Adding a black box in there seems incredibly misguided, no?
The same way I'd not trust an SRE when they've "got a feeling", I don't intend to start trusting Jev. The audit of process doesn't allow audit of reasoning.
yimingsu 10 hours ago [-]
Auditability does not come from an LLM or Jev but from the human operator themselves. Our Jev-driven diagnosis pipeline here is in the same role as all the AI SRE technologies on the market: they are tools to the human operator. If the SRE simply trusts all the AI tools, then I also wouldn't trust them just as when "they've 'got a feeling'".
codeduck 10 hours ago [-]
Your pipeline is then effectively nothing but an expensive flowchart.
yehosef 18 hours ago [-]
I had the same feeling. it seems like jev is a classifier that's smarter than bag of words but I would not think it would do well to ask it to judge code fix effectiveness.
soltanov 23 hours ago [-]
Premise is flawed: SRE is not a throughput problem, it is an accuracy problem.
sdcfgy 20 hours ago [-]
Yes and with an incredibly large context and incredibly large domain specific knowledge.
We had an LLM based SRE product from a vendor I won't mention because we had an NDA as our management is sucking them off by trading whitepapers for discounts. It was like a drunk monkey with a wrecking ball. Think we had to pull it in under 2 weeks because it took out multiple production systems and lead to an entire cluster failover.
The meat sacks now know they have job security.
PunchyHamster 19 hours ago [-]
My experience is that SRE-esque work you experience the full range of model, from "wow, it would take us ages to even get to this theory for why something failed" to "well the priority field in this protocol goes from lowest=most important, but LLM wrote code as if it was highest=most important, so if someone actually pushed it to production the company wouldn't have working internet access any more".
And a lot of "wishing up a feature", I was using it to explore solutions for a given problem in too I didn't knew 100% and it pretty much came to same solution I wanted to do but... the capabilities were not there in the tool so it just started making up probable config clauses, and of course, it didn't work.
Even on simpler stuff there were traps, for example in middle of debug session I asked it to modify Gitlab config to add request duration logging, so it added correct config format to a flag that didn't exist (option was there, just under different name), because it didn't bother to read the docs (since then I generally link it the docs first so it doesn't try to remember and get it wrong).
All of that is both very dangerous, and also easily fixed by just having competent operator there. And as a tool it's great, as replacement it is just AI bros delusion
jasonjmcghee 22 hours ago [-]
I'm guessing you're getting downvoted due to dismissive tone, but I think there's truth to this.
Labeling something as "don't escalate" that should have been isn't great. Paying 5x and having to wait 10s instead of 100ms or whatever for a reasoning model (with tools?) is very likely worth it.
yimingsu 10 hours ago [-]
We want something fast so that it can help the human SRE. Jev won't be deciding what to escalate or de-escalate anyways. Preferably, both the human and LLM will be pinged at the same time and work in parallel.
olgava 21 hours ago [-]
> SRE work is probably best suited for very few critical decisions
Yeah, and the article puts the cost of all 105 diagnoses at about $0.15 in Jev calls. At a dozen pages a day, I wouldn't worry much about the bill for either model. I'd be more interested in the pass rate: 76.2% vs 77.8% for GPT-5.6 Sol (medium).
I can see trying a cheaper model first if you're handling lots of requests and it can resolve most of them without escalating. A dozen pages a day doesn't seem like a reason to add that extra step.
yimingsu 12 hours ago [-]
Yeah, if the volume is small, then cost would not matter as much. But like I mentioned above, faster analysis and lower cost (than a full-fledged LLM!) is the goal, especially considering how much faster Jev is than an LLM. A Jev-driven pipeline can be a quick fast-path filter for a lot of alerts that are recurring or even false-positives.
endangeredhuman 22 hours ago [-]
This is an interesting experiment. Curious why you used LLM-as-a-Judge (gpt-6-astra). Did we have a ground truth of actual RCA done by a human to compare against?
yimingsu 13 hours ago [-]
Yes! All faults in SREGym-lite (and the parent benchmark SREGym) have their ground truths reviewed by us to ensure that they are accurate to the incident. In addition, we also conducted an analysis to ensure the LLM judge agrees with the human annotators too, which can be found on page 6 here: https://arxiv.org/abs/2605.07161.
N_Lens 23 hours ago [-]
Lots of hype-mongering around Jev atm. Color me sceptical.
Why? What’s the end goal? Lower cost? Faster analysis?
I configured an agent to respond to pages via Slack. It has access to ClickStack (for telemetry), Kubernetes, and GitHub. It runs one of the Sonnet models. The cost is so low, responding to less than a dozen pages a day (mostly from a very sensitive error count alarm) that swapping got Jev makes zero sense. This is especially true if the results are less trustworthy.
We can also add a confidence output to this pipeline, so we invoke a heavier reasoning agent (like Sonnet) when something suspicious deserves a closer look.
An SRE goes and fixes the cause of the recurring alerts.
And to answer the obvious question: because latency. This thing turns on the light or water valve in under 100ms on a laptop from 2001 with 8gb ram .. all running locally on the laptop.
The same way I'd not trust an SRE when they've "got a feeling", I don't intend to start trusting Jev. The audit of process doesn't allow audit of reasoning.
We had an LLM based SRE product from a vendor I won't mention because we had an NDA as our management is sucking them off by trading whitepapers for discounts. It was like a drunk monkey with a wrecking ball. Think we had to pull it in under 2 weeks because it took out multiple production systems and lead to an entire cluster failover.
The meat sacks now know they have job security.
And a lot of "wishing up a feature", I was using it to explore solutions for a given problem in too I didn't knew 100% and it pretty much came to same solution I wanted to do but... the capabilities were not there in the tool so it just started making up probable config clauses, and of course, it didn't work.
Even on simpler stuff there were traps, for example in middle of debug session I asked it to modify Gitlab config to add request duration logging, so it added correct config format to a flag that didn't exist (option was there, just under different name), because it didn't bother to read the docs (since then I generally link it the docs first so it doesn't try to remember and get it wrong).
All of that is both very dangerous, and also easily fixed by just having competent operator there. And as a tool it's great, as replacement it is just AI bros delusion
Labeling something as "don't escalate" that should have been isn't great. Paying 5x and having to wait 10s instead of 100ms or whatever for a reasoning model (with tools?) is very likely worth it.
Yeah, and the article puts the cost of all 105 diagnoses at about $0.15 in Jev calls. At a dozen pages a day, I wouldn't worry much about the bill for either model. I'd be more interested in the pass rate: 76.2% vs 77.8% for GPT-5.6 Sol (medium).
I can see trying a cheaper model first if you're handling lots of requests and it can resolve most of them without escalating. A dozen pages a day doesn't seem like a reason to add that extra step.