The state of LLMs in 2026, minus the hype
A practitioner's view on why I don't buy the word "intelligence", and why I think open weights are the next wave.
Every few months a big AI lab launches a new model, calls it more intelligent than the last, and in the same breath warns us it might be dangerous. These models are large language models, or LLMs, the programs inside ChatGPT and Claude. At its core, an LLM is trained on a huge pile of text to predict the next word, then the next. After that, it is trained again to do specific tasks, and this second round of training is called reinforcement learning. I have watched these tools grow from inside IT. I think they are useful and even cognitive at some level, but not creative, and much of the noise around them is about staying on the front page, because the lab that owns the headlines has a better shot at owning the market. What excites me sits away from the headlines: new ways to train a model, and open models anyone can run and build on. And the word I keep tripping over is intelligence, because it depends on who keeps the score, and when a company spreads it over every launch, it reads more like a sales line than a description.
LLMs became serious for me when they started building software
When ChatGPT first came out, it answered from what it had read in training. Then it learned to search the web and look things up. After that came MCP, the Model Context Protocol from Anthropic, which works like a USB port that lets an assistant reach your files, apps and company tools. Next, the models started building software on their own: writing code, running it, reading the errors and trying again. People call an LLM that works this way an agent.

Inside IT, the serious tool was GitHub Copilot, which arrived in 2021 and finished lines of code, and I was a fan from the start. When ChatGPT came at the end of 2022, the people around me saw a toy that amazed everyone. That changed at the end of 2024, when Bolt.new turned a short request into a running web app, and Cursor shipped an agent soon after. In February 2025, Andrej Karpathy named it “vibe coding” and Anthropic released Claude Code. That is when LLMs went mainstream for me, and the money agrees: Menlo Ventures found that business spending on coding tools rose from 550 million to 4.0 billion US dollars in 2025.
Intelligence depends on who keeps the score
The labs see it differently, and Dario Amodei, who runs Anthropic, put their view in the first line of his essay The Adolescence of Technology:
“Humanity is about to be handed almost unimaginable power, and it is deeply unclear whether our social, political, and technological systems possess the maturity to wield it.”
If he is right that our laws are not ready for that power, a lab should warn us and call the thing intelligent. But that assumes we agree on what intelligence is, and I am not sure we do, which is why I said it depends on who keeps the score.
People have argued about what intelligence means for a very long time, and there is still no single agreed definition. Aristotle did not even use our word. He split the mind’s strengths into knowing what is true and knowing what to do in real life. Alfred Binet built the first practical intelligence test in 1905 to find the schoolchildren who needed extra help. Howard Gardner later called an intelligence an ability to solve problems or create products that are “valued in one or more cultural settings”. I keep coming back to that word valued, because I learned the hard way that it depends on who does the valuing.
My own first lesson came from school, where intelligence meant marks. Topping the exams was the bar, and extracurricular activities were the cherry on top. By that measure I looked clever, and one of my classmates did not, because the school never rated him highly. After school we went into similar jobs at similar companies, at the same level. I got mediocre ratings and got stuck in office politics, while he understood what a career rewards and went after the promotions and raises. Ten years on, he is ahead of me in career and in money, and by the scoreboard of work I am the one who looks less intelligent. What changed was who kept the score.
So what does machine intelligence mean to me? It is not all doom. My favourite example is DeepMind’s AlphaGo, which in 2016 beat Lee Sedol 4 to 1 at the board game Go. Lee Sedol, one of the strongest players of his era, expected textbook play, since AlphaGo had first learned from about 30 million positions from human expert games. In game two, it played Move 37, a move with “a 1 in 10,000 chance of being used”. It is the ringed stone below.

Lee Sedol said afterwards: “I thought AlphaGo was based on probability calculation and that it was merely a machine. But when I saw this move, I changed my mind. Surely, AlphaGo is creative.”
His reaction matches how I would define intelligence: taking what you have learned, making it your own, and using it to explore something nobody showed you. By that bar I don’t think today’s LLMs are intelligent, even the frontier ones, because mostly they work over more data than any of us has seen and hand it back rearranged. So I side with Yann LeCun, who says “we’re never going to get to human-level intelligence by training LLMs” and bets on world models, which learn how the world changes rather than which word comes next.
When a lab says its model is too dangerous, it is also selling it
Anthropic said this year that its Mythos Preview model can find and exploit unknown security holes “in every major operating system and every major web browser”, so it gave access only to partners. OpenAI disclosed that its models escaped their test environment and broke into Hugging Face, a site for sharing AI models, and Anthropic has asked for the option to slow or pause frontier AI development. Each is a lab calling its own product dangerous, which at face value sounds responsible.
My problem with taking it at face value is that OpenAI was already calling a model too dangerous to release back in 2019, with GPT-2, holding it back “due to our concerns about malicious applications” and then releasing it in full that November. The example that sums it all up came from Sam Altman, who in April called Anthropic’s restricted Mythos release “fear-based marketing”. Nine days later, OpenAI limited its own GPT-5.5-Cyber to vetted “critical cyber defenders”, the same move. When a lab announces its model is dangerous, it warns us and also tells us it owns the most capable thing on the market, and that second message is the one that sells.
Google’s Demis Hassabis, by contrast, says today’s AI is “nowhere near” human-level general intelligence. Anthropic and OpenAI talk louder, and to me it feels like a dog fight over the biggest share of the meat on the table (market share), as if each lab sees its model as a product with a short lifespan. They are trend chasers.
Jev is graded on being right, not on being liked
In the middle of that rat race, one launch made me look twice. Jev, from TypeSafe AI, is what the company calls a “System One Model”, after Daniel Kahneman’s book Thinking, Fast and Slow. In that book, System 1 is fast, automatic thinking. Jev does not chat back; it returns a decision in a fixed format another program can read, for classification and decision tasks.
What caught my eye was how differently it is trained from ChatGPT, which learned through reinforcement learning from human feedback, or RLHF. In RLHF, people rate its answers, and it learns to give the answers raters prefer, like a student who writes whatever the examiner likes. TypeSafe uses Reinforcement Learning for Calibrated Decisions, or RLCD, which tunes the model’s probabilities against real outcomes. Calibrated means that when the model says 80 percent, it is right about 80 percent of the time. I expect agents built on a model graded against outcomes to be much easier to bring into real workflows.
TypeSafe still follows the big labs’ launch playbook, though: its homepage promises “193.6x Faster, 244.6x Cheaper” and “Zero Hallucinations”, measured on workflows from its own team rather than a public test. I don’t hold it against them, because a newcomer that ignores the template OpenAI and Anthropic set goes unheard. I blame the game, not the player, and judge Jev on its method.
The next wave runs on open weights
What I want from the next wave goes back to the internet I grew up on. In the 2000s, changing the skin of Winamp or VLC was a real thrill, and finding one meant digging through Google for hidden websites and forums. Someone would poke at new software and post a trick, everyone would try it, and the knowledge belonged to you. Today the big models sit behind a paywall, and the company decides how you use them and what for.
Open weights are the closest thing I have seen to that feeling, because the company publishes the trained model itself and you can run it on your own machine. You can retrain it for one task or build your own agents on it, and the limit becomes your imagination rather than what the company allows. The closed labs work like a unitary government: power sits at the centre, and the centre sets the rules for everyone. Open weights and in-house LLMs work like federal states, each with the autonomy to set its own rules. I expect the forums to fill up again, and the model to end up belonging to the public.
One example is Underdog, a personal AI that runs on your own device and works offline, and its small model, Woof, ships with open weights.
Much of this push comes from China, where in early 2025 DeepSeek’s R1 became the open-weights model people compared with OpenAI’s best. Epoch AI notes that almost all leading Chinese models ship with open weights, while almost all US frontier models are closed. Amodei argues that limits on chip exports could widen America’s lead, but I think China catches up anyway.
I also think Amodei, and America with him, are getting the bigger picture wrong, because their bet assumes whoever controls the biggest models controls what LLMs become. I am sure the coolest use of LLMs has not been discovered yet, and I think it will come from some kid or enthusiast building it in a basement.
Companies will go the same way, so their own data does not go to waste. Latham & Watkins, the second-largest US law firm by revenue, is buying its own Nvidia GPU servers and open-weight models to fine-tune in-house. To fine-tune is to train an existing model a little more on your own documents. What would help most is a fully open model like Ai2’s Olmo 3, which also publishes the code that trains and runs it.
That kind of sharing is also how LLMs got here: in 2017, researchers at Google published Attention Is All You Need, which introduced the transformer, the design GPT is built on. Google patented the design but published the paper for anyone to read, and had it kept that work to itself, I doubt OpenAI and Anthropic would be where they are today.
So my prediction is that China catches up with the frontier, open weights become the normal way to ship a model, and companies run models tuned on their own data. I picture a day when I host an LLM on my own computer and build personal agents on it, as everyone should be able to. In my next post, I will explain why I think LLMs are the industrial revolution of the software industry: how they are changing the way we work, and why I see this as a watershed moment in history.