Finally, a model showed up that's stupidly cheap at analytical grunt work — and it's honestly starting to look like it's coming for our jobs.
Jeff Nelson ran the experiment I've been waiting for: Jev against BigQuery's own AI.IF and AI.CLASSIFY, same tasks, scaled from a single row all the way up to ten million. Not a demo — a real side-by-side on public datasets, with the receipts.
The quality part is almost boring. Sentiment classification landed within a point across all three approaches. Topic classification split by 2.6 points, mostly because Jev kept filing tech articles under "Business" instead of "Sci/Tech." Fine, whatever — one fuzzy category, forever.
The cost curves are where it gets uncomfortable. BigQuery's standard mode scales with every row: $4.10 for 100,000 articles. BigQuery's own "optimized" mode barely moves — about $0.077 whether you feed it 7,600 rows or 100,000, because it trains a tiny model on embeddings after Gemini labels a small sample, then just reuses it. At ten million Stack Overflow questions, that same optimized mode finished in under two minutes, at 94% agreement with the actual tags, for basically the same price per run as the small test.
Then came the part that should make anyone doing this for a living sit up: stack Jev on top, using its own confidence score to escalate only the rows it wasn't sure about. Only 14% of articles needed the expensive model. Same accuracy as running the expensive model on everything. A third of the price.
Nobody needs a frontier model babysitting every row. You need a cheap one filtering for the 14% that's actually hard, and something bigger picking up just those. The "read this, classify that, flag the weird ones" tier of analytics just got a real, working blueprint for costing a fraction of what it used to.
So yeah. Finally, a model showed up that's dirt cheap at analytical work. And it's starting to look like it's coming for our jobs.
From Zero To Offer - FREE Product Analyst Playbook - get it here ignatenko.gumroad.com/l/wylyes



