Where Anthropic f'ed up was treating their monetization the way they treat model training. Turns out that success in experimentation is not transferrable.
They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling:
"You can only use Fable for a week as a part of your plan"
"Be ready! You have to start paying per token!"
"Nevermind! we extended it for a couple more weeks"
"Wait, now it's up to half your usage"
"Ok, now its..."
Most people want to not care. We want our AI like electricity -- Kind of just there no matter how easy/hard is for the supply. You don't want your electricity company to be on the brink of cutting you off any second.
That's Anthropic. You don't feel they want to give you a dependable service for an, albeit premium, price. It's a constant bargaining game. That forces people to look beyond the walled garden. There, they find models that are fine... and without the shenanigans.
fluidcruft 23 hours ago [-]
Yeah, I agree with this. The constant state of "...will the rug be pulled?!?" does discourage relying on it as a model and building a workflow on it. Anthropic used to just be a reliable thing you could play with. Now it's this constant source of anxiety.
It also didn't help that the government yanked it which adds another source of anxiety since OpenAI is on much better terms with the administration and the administration seems corrupt enough that they would mess with Anthropic if they got a big enough donation from OpenAI.
But anyway after Sol entered the picture, I don't think Anthropic can get away with this as much and I also think they're going to face a massive backlash from Max subscribers if they do end up ending the +50% promotion at the end of the month because Sol is a Fable peer and priced very competitively.
cle 20 hours ago [-]
My wife's startup made the mistake of building her internal operations around Claude Team.
Then she hired a VA in the Philippines. Anthropic promptly banned her account without warning once the VA connected to the account. It took her weeks to get her account reinstated, at which point she had already moved on to OpenAI.
kbrannigan 19 hours ago [-]
How many big tech companies let you talk to a human to get support.
Automation is wonderful to cut cost for them but for the users being unable to get support is a horrible experience.
But you cannot go elsewhere because they are the only player in town.
How can small companies with 1000x less money able to provide live support, but if you pay 20, 100, 200 dollars for a subscription you dont have a phone number to call ?
DanielHB 15 hours ago [-]
I think about 10% of my SaaS company is in the support department, no outsourced support at all. We have about 50k paying customers.
You message support, some real person reads and gets back to you within a day.
aashu_dwivedi 14 hours ago [-]
> I think about 10% of my SaaS company is in the support department, no outsourced support at all. We have about 50k paying customers.
What your your company do? Is it low ticket business or a high ticket business?
DanielHB 7 hours ago [-]
I work with music streaming, I don't really know how to judge what is low-ticket vs high-ticket. However it is a medium-to-high margin business.
Our free tier is time-limited but we still look at all tickets even from non-paying customers (in the hopes of converting them). A 1-hour intervention from a customer rep can result in a multi-year paying customer.
abirch 7 hours ago [-]
Vanguard does get some support right as they tier their support based on how much you're worth. Businesses need to focus on Lifetime Value of their customers and realize that some of their marketing budget would be better spent in support.
wing-_-nuts 9 hours ago [-]
>Automation is wonderful to cut cost for them but for the users being unable to get support is a horrible experience. But you cannot go elsewhere because they are the only player in town.
I can't tell you how many times I've experienced this with comcast. The last time I had to deal with it, was when I bought a new cable modem. I call in to provision it, the automated system assumes I have one of their modems and fails. For some reason I can't get technical support on the line and finally I resort to yelling 'cancel my account' over and over again until I finally get someone on the phone.
The guy was able to solve the issue in 5 minutes flat. The problem with automation is it's only ever going to be able to handle the 'happy path'
connicpu 8 hours ago [-]
It feels so degrading talking to the bot. Last time I did it I was trying to upgrade my service to take advantage of a 2.5GbE modem I bought and I almost said screw it because it was so frustrating with the long pauses after everything I said!
elictronic 8 hours ago [-]
The company solves the happy path every time. Your problem is they are solving their happy path which is profit optimization. The system is not poorly designed, it is working as intended.
The solution here is removing corporate monopolies and political power.
kevin_thibedeau 4 hours ago [-]
Threaten to sue the company. The bots will connect you to a human lickety split.
esseph 9 hours ago [-]
> How can small companies with 1000x less money able to provide live support, but if you pay 20, 100, 200 dollars for a subscription you dont have a phone number to call ?
They spend the money which can drastically cut into their profits.
rhdunn 13 hours ago [-]
Part of it is scale. If you have a small number of clients/users it is possible to provide that support. If you have millions or billions of users then you can't scale the support to handle the support requests, so some form of automation becomes inevitable.
exceptione 11 hours ago [-]
If you scale up customers, you scale up support. If you don't want to serve new customers, tell them to take their business elsewhere. But if you want to be a big boy, you need to play like you are one.
chasd00 7 hours ago [-]
> you scale up support
you'd end up with a call center larger than most cities. It's not feasible.
exceptione 5 hours ago [-]
If you think you hit a limit where you can't support any new customer anymore, you just tell them you have no capacity at the moment. Like any normal practice does.
Majromax 6 hours ago [-]
If your maximum addressable market is “the whole economy,” as seen in SpaceX filings, then a city-sized call centre (distributed, of course) really is ‘t that much of an ask.
esseph 9 hours ago [-]
But how else can they get an edge over any possible competition so that they can grow faster? Quarterly reports are coming faster and faster!
catlifeonmars 10 hours ago [-]
So you overextend, such that the quality of your support suffers? You can just call it what it is: greed. Have you considered that maybe a company shouldn’t have millions or billions of users? That’s a lot of eggs to put in one basket.
rhdunn 7 hours ago [-]
You mean like Microsoft, Google (GMail, YouTube, Android), Apple, Facebook, and others?
Sorry, you can't buy an iPhone because Apple has got too many customers.
Sorry, you can't have a GMail account because Google has too many customers.
intended 12 hours ago [-]
That's not scale, Its profit margins.
If firms had a base degree of customer support they were expected to provide, they would still exist. They would just not be as profitable, but customers would be better off.
I seem to remember there was a time when S/W was also designed with the aim to be easy to use, so that the need for support was reduced. It feels like the lesson learned was to keep costs low, not to ensure users were ok.
jgalt212 12 hours ago [-]
It depends more on revenues per customer.
WhyComboNadir 7 hours ago [-]
This is why I re-did the AI operating system I originally placed in Claude. For my 2.0 version I pulled it into OpenClaw (then eventually migrated to Hermes). All our work can be preserved after we change the model, even at a moment's notice or temporarily.
Right now we're using OpenAI's models by default since (unlike Anthropic) will allow us to use our pro subscription rather than token metering, but I've already had the joy of being able to change it to Kimi K3 (via OpenRouter) for an hour to try it out, and there were zero hiccups.
mv4 6 hours ago [-]
Hermes is the way to go. Accounts provided by the frontier labs are just too volatile.
osigurdson 16 hours ago [-]
A lot of companies seem to want to lock in to one solution or the other - like picking Oracle or SQL Server. The landscape is far too unsettled for that imo.
javier2 7 hours ago [-]
we have 7 coworkers we have been trying to re-instate for nearly 4 months now. All using the same google workspace sso, so there really was no special reason to ban them...
stuaxo 8 hours ago [-]
Misread as "infernal operations" which made it fun.
nullocator 20 hours ago [-]
But Philippines is on Anthropics list of allowed countries?
wahnfrieden 19 hours ago [-]
It was flagged for account-sharing, or detecting a compromised account
rhdunn 12 hours ago [-]
I had my GMail account locked for 1-2 years because I accessed it from my parents house (in the same country but in a different county) while on holiday. That was because they detected the account being used from a different IP address.
Using VPNs can also trip this.
willmadden 9 hours ago [-]
Interesting, I only use VPNs and have never had an issue other than constant "prove that it is you" secondary authentication.
Grimburger 19 hours ago [-]
How do you know there wasn't account sharing though?
I've seen some amazingly dodgy stuff when hiring people from south east Asia, sharing a paid account with friends worth a months rent there seems milquetoast in comparison.
coldtea 14 hours ago [-]
>How do you know there wasn't account sharing though?
Even if it was, it should be able to be sorted, maybe pay some overcharge or explain, and have your fucking business access re-instated.
anon373839 13 hours ago [-]
This is such a tough problem. Anthropic would need access to some kind of technology that could, like, intelligently handle unforeseen circumstances and nuances. Yeah, that’s definitely not something we should expect of them.
antonvs 12 hours ago [-]
It’s interesting how much Anthropic itself is a demonstration that its own hype is false.
serial_dev 11 hours ago [-]
Just use any product going all in on AI hype. In 5 minutes you will see annoying bugs, server is down, non-sensical press releases, confusing UI. This is GitHub, this is Cursor, this is Anthropic, this is Google, this is all of them.
anon373839 29 minutes ago [-]
Agreed. I think LLMs are best used as pair programmers or typists for users who already know what they’re doing. Or as tutors for users who want to learn.
Vibe coding is mostly garbage. But it can be useful for creating instant, disposable prototypes to investigate an idea or design direction.
coldtea 10 hours ago [-]
Including most of vide coded apps one sees, even from people who they'd trust before.
Anecdotal example, I downloaded a new alerting app recently from an indie dev who had a small following back in the day in iOS space. It asked for a subscription, like $20/year.
I thought, let me try this the (final version, from Mac App Store) app first. Well, it's a barely-there vibecoded shit. There's a bare-bones list, everything looks like my nephew designed it, the macOS "app" is a iPhone-size view of the iOS one, it has a bug that if you click on it it opens multiple duplicates of the same list view for no reason that you have to manually close, and in general it barely works.
Yay for vibe coding.
Forgeties79 11 hours ago [-]
Not to mention if their tools were so clearly useful they wouldn’t spend so much time making UI updates designed to force, trick, or confuse me into using their tool when it wasn’t my intention. SaaS companies with assistant integrations are the worst about this (looking at you, HubSpot)
svachalek 18 hours ago [-]
There clearly was account sharing, someone logged into the wife's account from the Phillipines.
jurgenburgen 17 hours ago [-]
The core complaint was that it took weeks for Anthropic support to restore access. For a startup that might as well be years.
weird-eye-issue 17 hours ago [-]
You aren't supposed to add other users to your Team plan?
17 hours ago [-]
techpression 16 hours ago [-]
Well a VA would need access to her account, just like how they would need email access etc.
But Anthropic has clearly picked the enterprise side of things, small teams and startups without millions of dollars in token budgets are irrelevant to them. I’m surprised they even got their account back to be honest.
thaanpaa 15 hours ago [-]
And yet enterprises will be the first to move to on-prem LLMs as soon as they become feasible. The next generation of TPU chips already promises 5x efficiency and enough RAM to run a 1TB+ model, and Kimi K3 is about as good as Fable for a lot of tasks, so we'll be there much sooner than anyone anticipated.
techpression 10 hours ago [-]
One can only hope, it would be great for everyone, myself included (the company I work for rather)
PunchyHamster 14 hours ago [-]
why does that matter ? Why does that block entire corporate account not a given user ? Use your brain
nullsanity 20 hours ago [-]
[dead]
adriand 23 hours ago [-]
As soon as I started using Fable I was like, okay, this is probably as good a model as I will need for software engineering going forward. I still feel that way. I don’t need a better model, I need a faster Fable.
The thing I miss most about programming is flow, and the constant bouncing between terminal tabs sucks. I’d love to do one thing at a time, with Fable, quickly.
kbrannigan 19 hours ago [-]
remember 4 year ago we use to : have stack overflow open, documentation, obscure forums plus other tabs.
An ide open with 20 tabs open each file a component, a class or an interface
We also use to hold entire codebases in our brain.
Zylokloto 13 hours ago [-]
Yeah StackOverflow which was either telling you to use google or it was so specific, that no one wanted/could respond.
Even a year ago when i was trying to do a hugo template manually with the help of the documentatin /tutorial, it was shit. The LLM at that time, was better helping me than the documentation.
kbrannigan 10 hours ago [-]
But it worked and it was so useful that LLM companies siphoned their data. It was how i learned programming
Zylokloto 8 hours ago [-]
I personally do not remember this time of Stack Overflow.
I had some helpful people helping me on IRC / Quakenet.
But the hugo example i found very interesting because it was the latest hugo ducumentation and I don't think I was able to find a tutorial. I tried it without an LLM first.
whateveracct 17 hours ago [-]
i still do :)
embedding-shape 16 hours ago [-]
Right, instead of having 20+ tmux panes with docs, specs and whatever, I just have 20+ tmux panes with various Codex sessions for various purposes instead, some of them been idling for days now, waiting for me to come back.
Things just moved up on the abstraction-ladder, but it's still there, hidden beneath all the TUI sessions instead.
echelon 17 hours ago [-]
And now I'm getting 10-20x as much done. I'd say the trade off is worth it.
I'm struggling to scale myself even further. This tech is unreal and I have so many things I can do.
For the first time, tech feels like the 90's-00's again. Everything is greenfield and exciting and big tech is struggling to figure out what to do about it.
People are just hacking all kinds of stuff, and it's awesome. Feels like techno utopia.
lelanthran 12 hours ago [-]
> For the first time, tech feels like the 90's-00's again.
It feels like the opposite of 90s - 00s: they were filled with periods where a person could self-study technology and get a job using those skills that few others had.
Where we are going (according to the AI-proponents) is children being able to replace you.
In brief; the 90s - 00s were a skill-valuation time, now we are looking at a skill devaluation time.
Unless you meant to say "Just like how any kid who could write broken HTML t put up a webpage could pretend to be a skilled professional, that's where we are now"...
kbrannigan 9 hours ago [-]
We used to get paid to code now we think we need to pay a subscription fee just to write software. They push marketing campaigns saying Manual coding no more, just vibe code, gain 10x speed for $100.
Then they hit you with hourly and weekly limits, you are wondering when you are going to get cut off. Since LLM at probabilistic it often feels like pulling the lever of a slot machine, hoping our prompt is the jackpot.
To make sure we win, we come up with systems, convoluted agents, context pipelines, rags to load . It feels like it's working, then bam you reach weekly limits.
(Just pay more if you want to keep winning).
I think developers need to wake up.
I myself started to use AI like a fancy debugger ,explainer. I make it walk me though every single line of code it writes.
I notice that i run into limits less, if i get cutoff, i can still make changes
chermi 6 hours ago [-]
As we've clearly seen, technology is static and the price to performance will never go down. And never be infiltrated by open source.
kbrannigan 5 hours ago [-]
I'm not trust I was meant to be sarcastic or if you're serious I cant tell Is it true?
noufalibrahim 16 hours ago [-]
I have mixed feelings.
The barrier and time between idea and usable implementation is almost zero now. I don't have to imagine. I can just write something and see it work before making larger decisions. I really like this. Many of my ideas were abandoned because I needed to study some obscure library. Now, I can learn the parts that I find interesting and just have the AI chew through the grunt parts easily. That's the good.
I started coding with a line editor on a small Casio handheld "computer" and used to keep programs in my head. I more or less knew what happened on each line without seeing the line. With larger programs, I had a mental model of what was going on where and a big part of the input to that was the effort of writing everything by hand. That's gone. It's not really important as far as the output of usable programs is concerned but there's a certain feeling of satisfaction that came with digesting a larger codebase and having it surrender it's secrets to you that's missing.
nananana9 10 hours ago [-]
I've seen this exact comment what feels like twice a day for the last 2 years, and not once have I seen the person making it back it up and show something even remotely impressive.
edgyquant 10 hours ago [-]
I dislike these comments just as much. What do you want people to show you? Most of us work on projects for other people where tasks that used to take a week take a day or less. It’s also a no true Scotsman as nothing we could show would be “good enough” because it’s just the same software engineering as before but faster.
But for instance I used to work in 1-2 client projects at a time and they take months now I can do 4-5 at once and they take a month. That’s a huge improvement
kbrannigan 9 hours ago [-]
I wonder if the project you build are more one off Disposable(sorry for my choice of words) software.
Do you have project that needs to be maintained. I am also interested in your workflow. Do you Vibe code , never look at the code or do you hold the LLM agent's hand.
I feel like it's a spectrum
edgyquant 8 hours ago [-]
I work on real projects with paying customers with the occasional mvp here, but I tend to select for people who have distribution so mvps become apps that need maintenance almost every time.
If anything maintenance is where it gets easier the mvp stage is where more focus is required
chasd00 7 hours ago [-]
> not once have I seen the person making it back it up and show something even remotely impressive.
heh this is funny because this reaction was all the rage in the 90s early 00s too. You'd put together something you thought was cool and then post a link on a forum only to be told how it wasn't even "remotely impressive". I'm glad people didn't give up back then and i hope no one gives up now.
monegator 16 hours ago [-]
BINGO!
bschwindHN 10 hours ago [-]
Software is pretty shit these days, and it doesn't seem to be getting any better. Definitely doesn't feel like a utopia to me.
mark_l_watson 11 hours ago [-]
I have lower standards than you: I pay for deepseek-v4-flash-0731 tokens from a fast and reliable US vendor and I feel like working on one task at a time is fast and gets almost everything done I need.
jmalicki 21 hours ago [-]
There is GPT 5.6 Sol on Cerebras if you want to try that experience for an ungodly sum of money (not getting into GPT 5.6 Sol vs Fable, but only one is available on Cerebras) for an 11x speedup.
hakimg 15 hours ago [-]
OpenAI also have 5.6 fast mode for a 2.5x speedup for 2x cost and is available on standard plans.
azalemeth 15 hours ago [-]
I _still_ haven't actually been able to _use_ Fable at all. Those safeguards just refuse biology in general.
jghn 11 hours ago [-]
It changed for the better a few weeks ago. Obviously it depends on what one is doing but I haven’t had an issue with my biology related material since that update
codazoda 21 hours ago [-]
You may or may not like agents mode. I also hate flipping tabs, but I enjoy using agent mode with well named sessions. I still stick with a single session until I must move to another, then I leave them around for a few days until I’m sure I won’t need to pick up where I left off again.
Command:
claude agents
pastel8739 21 hours ago [-]
I don’t think they were complaining about literally flipping tabs but rather just needing to context switch so often. This doesn’t sound like it helps with that.
Flow 14 hours ago [-]
I felt the same about GPT-5.2 on High. It did all I asked and it did it good and cheap. Too bad it’s no longer an option at all.
zelphirkalt 15 hours ago [-]
A faster Fable -- So you mean a model that's smaller, yet delivers the same quality of responses, ergo a better model?
osigurdson 16 hours ago [-]
I used Fable a bit when it was available. I didn't get the senae is was dramatically better than OpenAI. Now I am considering canceling my Claude sub since I can no longer experiment with their best models.
eastbound 12 hours ago [-]
> OpenAI is on much better terms with the administration and the administration seems corrupt enough to...
Which tells you everything you need to know about OpenAI.
20 hours ago [-]
visarga 18 hours ago [-]
> That forces people to look beyond the walled garden.
Every time I get "you used your quota, come back in 3 hours, or 2 days" -> that is experimentation time with their competition, leading to changed service plans. When they said "claude -p" will be billed at API pricing even for plan users I moved my harness off claude. After I integrated codex, then it was never going to be a full claude project again.
What business encourages users to try their competition and adapt their usage to the competing products?
zymhan 17 hours ago [-]
It drove me to setup Qwen 3.8 this weekend. I couldn't see the value in just giving them money for a higher tier plan instead.
I've never run a local LLM model before. Certainly won't take as long to iterate on this.
dexterlagan 16 hours ago [-]
Qwen 3.8 is excellent. With the right harness, it does about 95% of what Opus can do, in my case automation software development. Since 3.8 came out, I have significantly revised my expectations for a local model. Give it another year or two, and we'll be running fast and free local models for nearly everything that matters, and these costly subscriptions will be a thing of the past. I've always believed that AI should be free for everybody, like TV and radio. We're almost there.
abc123abc123 12 hours ago [-]
Free tv and radio? Where do you live? Where I live you either pay taxes for it, alternatively, it is so ad infested that it is not possible to watch it.
I suspect the same will/is happening with AI. Either you will pay for it, or it will be so ad infested that it will become useless.
mrtsepelev 15 hours ago [-]
What harness would you recommend? I’ve tried Pi but the model struggled to stay on track after the compaction.
I have only 48gb of ram, so can fit only 80k context max, so good compaction is must.
Scaled 12 hours ago [-]
Not op, but check out open code; you can turn on K/V quantization to help with increasing context if you have not already. I think K needs to stay at least 8 but I hear V can go down to 4?
wccrawford 13 hours ago [-]
I'd also love to hear your setup? How much VRAM/RAM, I assume Qwen 3.8 27b, what harness, are you using any particular skill set?
c16 9 hours ago [-]
I've a 32gb and 64gb (work) MBP. 32 works - just and sits at around 28/29gb of 32. 64 works great, so the 48gb laptop with MLX + MTP should be fine. I'm using Ollama.
I initially used the Claude Code harness on 3.6 A3B, but found that tooling would break as Claude released new versions and things would go weird. I've since written my own harness which has basic operations: read, find, bash (which can write files, python etc...) & web_fetch, all within a mac container. Works amazing. You don't need anything complicated to go very far.
Low hanging fruit would be Pi or OpenCode. If you really want a much better understanding of what your hardware is capable of then give writing your own a go.
Additional tip: Low Power mode reduces some token speed, but stops the laptop over heating and the fans going crazy.
ryreacher 7 hours ago [-]
What harness are you using for Qwen 3.8?
walthamstow 17 hours ago [-]
The claude -p thing was doubly stupid because they quietly allowed it again a couple of weeks later, so they pissed off developers for nothing
airspresso 15 hours ago [-]
Wait, it's allowed again? Completely missed that. Been avoiding to use it and trying to find workarounds, not great.
PufPufPuf 7 hours ago [-]
Yes, they sent an email about it. You can also use the Claude SDK with an OAuth token ("claude setup-token" output) and it counts against the regular limit. Maybe they were afraid of losing users dependent on ACP (Zed editor and other compatible tools), since Claude Code does not have native ACP support and integrates only through the SDK?
walthamstow 15 hours ago [-]
I can't find the page now but yes they quietly "paused" the June 15th rollout of API pricing for -p headless. Presumably to come back again one day.
smoe 7 hours ago [-]
I think they didn't even bother making a separate post about their backpedaling, they just slapped some disclaimers onto the existing page:
Update June 15: We're pausing the changes to Claude Agent SDK usage described below. For now, nothing has changed: Claude Agent SDK, claude -p, and third-party app usage still draw from your subscription's usage limits. The previously announced monthly credit, which would have been available to eligible claimants in connection with these changes, isn't available. We’re working to update the plan to better support how users build with Claude subscriptions. When we have an update, we'll share it before anything takes effect.
lelanthran 12 hours ago [-]
> What business encourages users to try their competition and adapt their usage to the competing products?
If you are selling something, and losing $10 on each sale, you also would want to limit how much you sell.
I mean, sure, you are losing money on each sale so you can landgrab, but you still have to balance the land-grabbing with how much money you can actually lose.
LaurensBER 17 hours ago [-]
>™Every time I get "you used your quota, come back in 3 hours, or 2 days" -> that is experimentation time with their competition, leading to changed service plans.
I guess this is why they're pushing Claude code hard (not supporting agents.md, not allowing third party harnesses, etc) but when switching to another provider is as easy as opening a new terminal and typing omp/pi/codex your moat is effectively zero.
They can compete on price, quality or value but anything else is just madness. Currently they (arguably) own quality but this won't last.
dan_ggggg 13 hours ago [-]
> What business encourages users to try their competition and adapt their usage to the competing products?
They are high on their own supply. The people running these companies are delusional imbeciles who have been placed in charge of billions of dollars.
neya 13 hours ago [-]
I'm surprised no one mentions about their recent privacy violation(s).
The breaking point for me was the privacy violation. They've been fingerprinting every request and violating users' privacy hoping no one would notice. Too bad, someone found out and that was the day when I cancelled my subscription.
I spend way too much time in all the LLM related subs, to the point that i consider it unhealthy (inc claude/anthropic subs).
Its in my opinion not wide spread at all and as today is literally the first time i ever hear anybody mention this.
neya 8 hours ago [-]
No, that was the very first time that article was submitted to HN. That's not called "talking about" it. Talking about it means highlighting this enough in discussions so users really know their privacy is being compromised. I have more respect for AI companies that openly talk about selling user data than the ones pretending to be privacy heroes while doing the opposite.
ryreacher 7 hours ago [-]
How this entire watermarking thing plays out will also be interesting
guluarte 9 hours ago [-]
Engineers love to play with different tools, in my company some use opencode,omp, hermes and you cannot use the team sub with those
Aurornis 19 hours ago [-]
> They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling:
The consumer side cheap monthly plans exist for the same reason companies like Cloudflare and Vercel have a free tier: When it’s cheap and easy to get developers familiar with the tools, they will push their companies to pay the real money for those tools.
It’s a hard balance with LLM serving because you can’t really make it free. $20/month is close to free, but the $200/month plans are in a difficult place where they’re big enough that many small companies pay for $200/month plans for their employees and ignore the enterprise features you get with the full expensive arrangements. So the companies are continually adjusting the $20-$200 plans to keep them from being reliable options for businesses, which is where the real money is.
There’s a short sighted cheering on of the 3rd tier and lower companies offering lower rates, but we’re already seeing them ratchet up the pricing and keep larger models closed after they get market attention.
cherryteastain 17 hours ago [-]
> There’s a short sighted cheering on of the 3rd tier and lower companies offering lower rates, but we’re already seeing them ratchet up the pricing and keep larger models closed after they get market attention.
Sorry, but people are cheering on Chinese companies (of whom your are unduly dismissive with your '3rd rate' comment given how good GLM-5.3, Kimi K3 are) not only because they are more economical, but also because they do not constantly refuse to do legitimate tasks and provide you with the weights for self hosting these models.
dv_dt 14 hours ago [-]
I feel like many VC driven companies have completely forgotten how to compete on basic value for product and instead tie themselves into knots with meta-competitiveness games.
TheOtherHobbes 14 hours ago [-]
"We have the smartest model in the world but our company consistently does stupid things" is not a sustainable business model.
Maybe Anthropic's enterprise sales are going brilliantly, and the rest of us are just pixel dust to them.
Still. Brand perception is a thing, and between rug-pull usage policies, weirding verbedly output quality, and "I'm sorry Dave I can't do that" pushback, Anthropic are clearly having strategy issues.
jaapz 11 hours ago [-]
Also, for whatever software engineering work I throw at Fable 5, Opus 5 also does fine. Apparently Fable is supposed to do better at long running tasks (in other words - burning more tokens without interacting with the user), but that's not the kind of work I'm doing.
After Fable 5 launched, it was better than Opus 4.8 for sure. Then they rug-pulled Fable from me (EU), and later released Opus 5. Now I only reach for fable when Opus 5 API returns 529 for the millionth time this year.
spaceywilly 10 hours ago [-]
Fable was really excellent before the whole fiasco with the US government. Once they brought it back it was not the same at all. I have switched over to ChatGPT now for most things, its answers are way better than Fable in my experience. I still use Opus 5 for purely coding tasks.
This is why I think open weight models will win out in the end. Right now there’s too much going on behind the scenes with the models. Day to day you never know if you’re going to get smart Claude or dumb Claude.
jayGlow 8 hours ago [-]
that part is extremely frustrating, I can't tell if it's a placebo or if the model quality does actually vary. the uncertainty makes me more hesitant to rely on it heavily as some days it just seems incredibly dumb to the point of being useless.
misalliance 3 hours ago [-]
For browser game generation, Fable 5 consistently produces much better game prompts and playable 2D or 3D prototypes than Opus 5 (based on 500+ prototypes I created using different models). It has a much better grasp of how visual elements work together and implementing game mechanics.
rob 22 hours ago [-]
As Anthropic does this, OpenAI Is giving everybody resets like every other day now on Twitter.
I'm strongly considering biting the bullet and just ditching my $200/month Claude Code plan for the Codex one instead, especially because I keep running into my weekly limits (even sticking to Opus.)
Petersipoi 20 hours ago [-]
I ditched Claude Code $200/month a couple of months ago in favor of Codex $200/month. The value is night and day.
1. No 5 hour usage limit
2. Weekly usage gets reset CONSTANTLY. It's crazy. The longest I've ever seen it go without a reset is maybe 5 days?
3. I don't feel like OpenAI is constantly trying to fuck with me. Unlike Anthropic. I would way rather have Sol all day every data, consistently, than a slightly better Fable for like, 1 prompt every 5 hours, and only when Anthropic decides to not treat me like a cyber criminal. Believe in yourself as much as Claude believes your CRUD app is going to hack the pentagon.
4. Getting access to image generation, though I don't use it too much, is a nice perk compared to Anthropic.
edit: Should mention that I had like 4 banked manual resets as well. It feels like OpenAI wants me to use their product, whereas Anthropic wants my money while giving me a nerfed experience
Aeolun 16 hours ago [-]
> The longest I've ever seen it go without a reset is maybe 5 days?
This last reset took 6’ish days. I know because I was almost out of limit.
pell 14 hours ago [-]
>I would way rather have Sol all day every data, consistently, than a slightly better Fable for like, 1 prompt every 5 hours, and only when Anthropic decides to not treat me like a cyber criminal.
I have used Fable heavily on the lower Max plan and you are really exaggerating the limits here. I've had many multi-hour sessions with Fable on Max.
Petersipoi 5 hours ago [-]
Maybe you use AI differently than me
pell 3 hours ago [-]
What are you doing that uses the entire Fable credits so quickly? Genuine question.
Revanche1367 20 hours ago [-]
I got the AI ultra plan for gemini, the models are lower quality than claude and codex for sure but it comes with a pretty high limit for my purposes and I haven’t reached the weekly limit yet after about 3 weeks of usage. I have enterprise claude at work, the budget isn’t particularly great and I keep running out within a week tops with any serious work. Not even considering it for a personal plan with how quickly the tokens run out for even the mid-tier model/effort combinations.
I’m sticking with gemini for now given the 20TB cloud storage and youtube premium that comes with it but I’m still open to switching to codex (which I’ve had a long term plus plan for). If google keeps delaying the pro models for much longer or makes them excessively expensive, I’m likely to switch out.
visarga 17 hours ago [-]
> I’m sticking with gemini for now given the 20TB cloud storage and youtube premium that comes with it
I used to have the Google One with Gemini and Drive space and YT Premium separately for my family. A credit card expiration lapse lead to closing YT Premium and being locked away. Why? it was because Google One plan bundles YT Premium Lite as an extra. So they actively blocked me from getting Premium back for a month.
Now I moved to YT Premium on my wife's account and downgraded One to lowest tier. Never heard of a company forbidding users to upgrade their plans before this. I suspect Google really wants users on Google One no matter what they want. I was barely using Gemini anyway, already have claude and codex plans.
DefamationStati 18 hours ago [-]
I've got the AI pro plan for Gemini and the limits as absolutely abysmal when using 3.7 flash in antigravity, If I use swarms then I can't even finish a single prompt without it reaching the limit.
technotony 20 hours ago [-]
Where do you get those? I got about 5 resets in July but none since then
wild_egg 20 hours ago [-]
You should have just got a reset today at least. There are several sites around for tracking them now. I use this one:
I got a reset today and also one of those reset tokens which I'll probably use sometime this week.
dannyw 23 hours ago [-]
I agree they’ve done a bit too many pricing / usage promotions and A/B tests.
The period was also marked with many billing bugs, like spending people’s usage credits for included Fable for a few hours (gave me a huge shock), but to their credit they refunded it.
hparadiz 22 hours ago [-]
They are fumbling the bag hard. AI's utility is for general purpose. The floor is rapidly improving from below them. With chatgpt I'm uploading all my day to day stuff. Meanwhile Claude is only for occasional super hard tech problems which are rapidly improving with being solvable easily by Sol. So what's the value add?
Their disrespect for their users is also another problem. You only get one shot to make a good impression.
8cvor6j844qw_d6 22 hours ago [-]
> With chatgpt I'm uploading all my day to day stuff.
Same, it's been a while since I logged into Claude web.
ChatGPT web usage being separate from Codex usage limit is a nice touch unlike Claude.
hparadiz 22 hours ago [-]
If I got cut off from asking chatgpt what the status of a spreadsheet is for handling my personal bills due to my rate limit from coding work I would not be coding on it at all. Glad someone over there realized this obvious fact.
mgkimsal 21 hours ago [-]
I'd go further and say chatgpt offering continuing service just with degraded model keeps me in their 'free web use' tier. I have API keys for coding stuff, but for most of my day to day use of public llm, I know I won't get 'blocked' using ChatGPT. Yes, the model may change, and I'll get 'worse' output, but usually I can't tell the difference. Using Claude for day to day stuff, I get completely locked out after so many hours. Again, I have API keys for Claude as well, and use it for 'pro' work, but day to day chat stuff... it's not my daily driver.
dexterlagan 16 hours ago [-]
They haven't communicated well the fact that the default model should be Sonnet 5, which should give you unlimited use for common coding tasks (say with occasional subagents use) on the Pro plan. Instead they're pushing Opus and even Fable, to try and get people addicted to the higher tier, without realizing that nearly everybody has a Sonnet for peanuts via DeepSeek V4 on OpenRouter, or completely free through Qwen 3.8 locally. I predict major trouble for Anthropic, now that OpenAI's models are closing in, are cheaper for daily use and don't have those silly 5 hours limits - and that Chinese models are getting really good and are even cheaper.
vikramkr 14 hours ago [-]
Sonnet 5 is a trash model and stupidly expensive if you accidentally set reasoning tokens high - more expensive than fable - it absolutely should not be the default lmao. If the common coding tasks you use ai for is doable with sonnet or local qwen - you're either not using Claude code (if you are, you'll very quickly see that sonnet 5 in Claude code is not a model for "occasional subagent use" - spinning up subagents is the only thing it's good at and it does it way too much. It can spin up subagents and waste huge amounts of tokens but it can't write good code lol.) or you've got Claude code workflow that is very human in the loop where you are significantly steering and controlling the models. And in that case your default should be to use gpt. Claude models are stupid slow.
datadrivenangel 7 hours ago [-]
Yeah Opus 5 on low is faster, cheaper, and better than Sonnet 5 on high...
praseodym 16 hours ago [-]
Claude Code also doesn't make it easy to make _efficient_ use of the different models to reduce overall cost. There are many tokenmaxxing features (e.g. ultracode that spawns dozens of subagents) to burn through the 5-hour limit in minutes, but if you want to let an Opus planning agent use Sonnet for implementing you have to orchestrate your own workflow. I'm pretty sure that's because the Anthropic employees working on Claude Code have unlimited token budgets so they're mostly on tokenmaxxing workflows themselves.
thisisit 14 hours ago [-]
Add to that at one point Claude was the go to models for the very basic use case for LLMs - text generation.
With newer models text generation outputs have gone from probably human readable text to dense philosophical treatise about "load bearing" and incomplete sentences. So much so that now you need skills or another LLM to just parse the output. Simple answers and text generation just doesn't exist.
coldtea 14 hours ago [-]
The problem isn't the confusion from pricing and ToS changes, but the constant feeling they don't give a fuck about their customers, and they'll fuck them up with lock-in tactics and high prices at any chance they get.
JohnMakin 7 hours ago [-]
Yes, this was a bit upsetting for me. I found myself organizing my work hours and availability around perceived or actual fable quota limits / trial periods - only to find out multiple times it didn't matter, there is no seeming strategy or rhyme or reason to it. Enormously frustrating, and I did end up just settling for a while with cheaper models.
I'm not saying this with any undue derision, it's genuine - do they have a real product team or are they clauding that too? The direction makes little sense.
sarjann 14 hours ago [-]
I think an important part is communication, so many of these issue could be fixed by saying "Hey we're seeing our utilisation go over x% over the weekdays so we need to implement "surge usage" during this period starting in 2 weeks.
Instead often it feels like they make a change, then wait for someone to figure it out. Then Anthropic ends up being reactive as opposed to proactive in communication.
It is kind of funny because surprises from OpenAI tends to be positive (Tibo resets), on the Anthropc side I dread them.
steveBK123 6 hours ago [-]
Isn't it just tipping the hand at where the actual businesses are going to end up inevitably?
The only B2C is going to be watered down ad-driven BS, and they will charge B2B via tokens.
The $50/mo - $200/mo consumer LLM subscription is not something I expect to last long / or to drive much of the revenue share... like individuals paying for Gmail vs Googles overall business.
snovv_crash 6 hours ago [-]
If they can make a business like that, yes. But with how rapidly the competition is catching up, I'm not sure that's a viable strategy.
sznio 16 hours ago [-]
In my case it's been kind of unhealthy, just constantly waiting for the token limit to reset, always feeling like I'm wasting a resource if I'm not using subscription right now.
I just put $30 on openrouter, switched to Pi, and I finally have a calm mind. Since I actually pay per request I want to maximize efficiency rather than utilization
AlwaysRock 8 hours ago [-]
Yup. It also seems like the latest and greatest is a smaller and smaller gap everytime. I will continue using the second best more affordable model until a new model comes out, everyone talks about how great it is, and the last greatest model becomes the second best and costs the same as the previous model I was using.
larodi 17 hours ago [-]
Opus and Sonnet 5 babble like crazy - this’ one reason. At some point one feels as if staring at the Random himself, not a conversation. Fable is super expensive.
From a cost-effective perspective the GPT models are much cheaper - one can easily tell it takes longer with GPT5.x to exhaust limits and this matters A LOT.
I can’t say which of these corpos I despise more though. I though for a while Dario was cool, but a massive distrust is piling and the first third player offering decent experience (and showing some decency) will win me over.
For the record - I’m also unsure whether I despise more Exxon or BP or burning fuel as a whole. Hope u get the point...
qaq 7 hours ago [-]
Not just that once I hit my Fable limit I naturally experimented with other options and realized that 5.6 Sol + Grok 4.6 gives me same quality of results
as Fable + 5.6 Sol so not really that reliant on Fable anymore.
ray_v 11 hours ago [-]
Not only that but it appears as if they're also treating the platform that customers actively use in such a way as well - heavily A/B testing features and behavior of the platform with little regard for customer comfort in terms of platform use, moving targets for subscription limits, etc etc. I suppose most of us chalk it up to the technology being new and evolving, but I'm not so sure everyone shares that same sentiment clearly.
lukan 12 hours ago [-]
Not just with the pricing models and avaiability, also with how the tool behaves. Currently I am fighting "auto-mode" that was introduced recently - and enabled by default without warning - and now I have to watch all the time that it does not got reenabled somehow again, depending on project and device I am developing it.
Also that the behavior of the models change, suddenly more fluff in the comments etc is annoying, but that is probably being part of using cutting edge tech.
fg137 11 hours ago [-]
FYI you can disable it globally
lukan 10 hours ago [-]
In theory yes, I know, but it somehow kept coming again (probably bugs, also I switch dev devices often). For now it seems off.
throwaw12 11 hours ago [-]
it also feels like they are optimizing their models to output more tokens, because everyone is saying inference is profitable, they wan't to close the gap between training and inference cost by increasing output token count (which also increases input tokens in agentic use cases), with 5 min TTL, this means you almost don't have a cache
pbreit 21 hours ago [-]
The incremental improvements seem like they are going to be pretty modest from this point.
usef- 21 hours ago [-]
For what it's worth, the things you describe are mostly because they're extremely short of GPUs and growth rates were absurdly high.
(Eg. They repeatedly said they'd keep fable in lower subscription plans if they had the capacity)
dkersten 12 hours ago [-]
I stopped using Claude because of this BS. If I pay for a service, I want to know what I’m paying for, I want it to be predictable.
Anthropic have been anything but. Flip flopping on model availability, model access behind an opaque filter, their past behaviour of model degradation as they prepared their next model… these are not signs of a reliable service.
I’ve mostly settled on using a mixture of open weights models through Together.ai and Fireworks.ai, a MiniMax subscription for high-token-use tasks that don’t need the best model (for $20 I get what feels like infinite tokens), and codex for the occasional high complexity task, although with Kimi K3 and hopefully soon GLM 5.3, it’s becoming increasingly less important. Deepseek 4 flash is my cheap main with delegation to other models as needed.
I’ve also found LFM2.5 8B surprisingly useful for single-focus tasks like “does this diff touch anything that isn’t related to the task”, and it’s incredibly cheap ($0.03/0.12 per M in/out).
kees99 11 hours ago [-]
I run LFM2.5 8B locally. It's not very smart, to say the least.
But then, there are plenty of mindless, menial tasks out there, and it would go through those like a champ. And quickly, too.
dkersten 7 hours ago [-]
Exactly, not all tasks require smarts, they just need to be good enough. I’ve found that if prompts are really focused, the tiny models do alright.
nsoonhui 16 hours ago [-]
Not entirely sure how OpenAI is any different. Their quota system seems random to me. I can use up my quota in a single day, and the next day it gets refilled for no apparent reason. But another time, I also used up my quota in a single day, and there was no refresh; I was made to wait six more days.
To me, all of these are just exercises in getting me to pay for more tokens at API rates.
SkyPuncher 21 hours ago [-]
Yea, I don’t have any interest in trying Fable because I’m not interested in the BS that’s going to come with it.
I’m at the point where I need stability and predictability. I want the B- student who shows up everyday rather than the A+ student that’s unreliable.
tmp10423288442 19 hours ago [-]
With Sol OpenAI is more like the A student, and Astra seems like it will be the S-tier student if the rumors are true.
AgentOrange1234 21 hours ago [-]
Yes, I feel like you can just sense the garbage coming. Age/identity verification, mandatory data sharing, morality policing, "Answer Engine Optimization" ads and influence peddling... it's going to be so painful to watch it all enshittify.
ryandrake 20 hours ago [-]
OP made the "electricity" analogy, and that's really all people want. I want to plug something into the wall and have it work. I don't want to have to worry that my electric company is going to rug-pull me because I plugged the wrong appliance in, or I didn't agree to some TOS, or I used the electricity to run grow lights for my pot farm, or this or that or the other.
z2 19 hours ago [-]
The analogy is apt because I feel this is exactly what keeps Anthropic and OpenAI's owners up at night -- becoming the utility company the People want them to be. Ironically their behavior may accelerate their fears. And yes, the so-called safety features are ridiculously invasive and the worst is agreeing to have surveillance cameras installed in every room that occasionally detect any attempt to grow plants with LED strips as a pot farm. After a false alarm of almost having my ChatGPT account terminated for cybersecurity abuse and appeals auto-denied twice (I did nothing even close to hacking), I have started doing everything I can to decrease switching costs and thus the bargaining power of the suppliers and I'm doing the same for my company.
ryandrake 8 hours ago [-]
Basically I want most companies to be selling basic, reliable commodities. I don't want their stupid value-add or lock-in. I don't want my electric company to sell me "MyElectric+, a subscription service that lets me (and them) enable and disable my appliances from the cloud and share my meter readings with friends." Just get the "electricity" part right, and I'm good. I don't want my water company to sell me "MyWater+, a subscription service that lets me (and them) flavor and carbonate my water with an advanced cloud service that...blah...blah...blah..." Just stop! Fire all your idea guys and just supply the base product. Hell, I'd be happy with a phone that no apps pre-installed. Look at phones today, so much built-in uninstallable crapware that it would make Gateway in the 90's jealous.
mrguyorama 8 hours ago [-]
People need to understand this took decades of regulation, lawsuits, and political action to make electricity as nice as it is.
zahirbmirza 11 hours ago [-]
There is a consumer unfriendly ethic behind this. Overly long answers are a cunning was to increase token cost and therefore profit. Consumers are savvy and will prohibit monopoly whilst there is still plentiful competition.
20 hours ago [-]
chrismsimpson 19 hours ago [-]
Also you don’t want to connect to the pipe and then after the fact find they’ve started diluting arsenic into it.
michaelbuckbee 13 hours ago [-]
It also distorts the testing as it encourages non-typical behavior.
itemize123 18 hours ago [-]
main reason is the 30days retention; not the plan changes
YetAnotherNick 22 hours ago [-]
I don't think Anthropic wants a stable experience on their consumer subscription plans. It is just used for customer acquisition who will then ask their employer to pay for enterprise plan(assuming most employer care about data control) which is based on tokens.
Most coders don't pay for tokens themselves. It's just on reddit and HN you would think that everybody does.
codebje 22 hours ago [-]
I pay for my tokens for my own projects, at least when the ones Google seems willing to keep throwing at me for free don't cut it. I'd think that's not too uncommon, especially here where there's likely a high ratio of hobby coders (whether also professionals or otherwise).
queenkjuul 11 hours ago [-]
Free tier Gemini is good enough for my personal projects and hopefully a local model can replace it before i ever personally pay for a single token from anyone.
Work can pay a Claude sub. At home i see no need
jayGlow 8 hours ago [-]
local models are getting pretty close. they still require some serious hardware to run at a reasonable rate but they're about smart enough to be useful.
YetAnotherNick 21 hours ago [-]
I also pay but just the $20 plan and just for chats/lightweight personal site editing. I get unlimited token usage from my company.
In my company the average claude token usage is something like $5k/month/employee. Most hobby coders don't spend anywhere close to it.
Vespasian 18 hours ago [-]
I strongly suspect most companies don't spend that much either.
Unlimited tokens aka unlimited cost is something that not every use case needs.
YetAnotherNick 17 hours ago [-]
I think most good engineering companies are spending >$1000/employee/month. Subscription users spend $100-200.
cactusplant7374 22 hours ago [-]
Dario said in an interview that they originally wanted to be an enterprise only company.
andai 22 hours ago [-]
What changed their mind?
gensym 21 hours ago [-]
I would guess that it's hard to compete with Google without a consumer component. Most enterprises already have contracts and policies with Google, so anything driven from the top-down is likely to prefer Google.
(And failing that, there was a real risk for OpenAI to be the default for enterprises.)
To defeat that, you need to frontline employees the chance to experience better tooling and models which is where the subsidized subscriptions come in.
trollbridge 20 hours ago [-]
Or AWS, or Microsoft, etc who have all the other stuff enterprises want.
etempleton 21 hours ago [-]
OpenAI is also now actively discounting their model. They just offered a free month to users. Both companies seem spooked by what I have to imagine is slowing user growth.
Zylokloto 13 hours ago [-]
Might just be preparation for their IPOs.
But the agentic layer is being worked on, agents will start consuming more and more tokens
anukin 20 hours ago [-]
Where is the free month offer going on? Are you talking about usage resets?
wavewrangler 19 hours ago [-]
he probably canceled his subscription and they offered him a free month as a result
Revanche1367 20 hours ago [-]
Some users on HN in recent months started describing Anthropic as having become a “token merchant” and I think that moniker is quite apt.
raincole 17 hours ago [-]
If they actually became a token merchant it'd be amazing. But they didn't. They tried to hide the chain of thoughts tokens. They banned accounts for using third-party harnesses with Claude subscription. Their tokens are also not very at a very competitive price.
yieldcrv 20 hours ago [-]
They need to fire their growth marketer
Their truth is “we don't have compute and are working to improve capacity”
People would root for that
Instead they got people rushing to escape the permanent underclass until they have a mental health crisis just to beat the fake deadline. $100, $200, is a lot for those people
ls612 23 hours ago [-]
The fact is that demand for tokens at electric bill rates so far outstrips what can be supplied currently not just with frontier models, but with open weights cheap models too. Running an always on Deepseek flash agent would cost three figures a month at API prices.
dannyw 23 hours ago [-]
Total costs sure, electricity only costs no. My two DGX Sparks run DS4 Flash at about 50tok/s concurrency=1 which is more than suitable; at about 150W total wall power when generating.
That’s about A$16 a month in electricity if I ran it 7x24x30.
troupo 15 hours ago [-]
> Wait, now it's up to half your usage
It doesn't help that all their models are bow trained to waste as many tokens as possible with their extremely verbose output
willmadden 9 hours ago [-]
The solution:
1) Set pricing tiers that do not change.
2) As models evolve, move the outdated models down the ladder, and replace the top tiers with the frontier models.
3) Give the users a warning before you do this, so they know their model is changing.
4) Sort the economic distortions out of your OPEX and reset pricing when the technology ossifies.
Is that so hard?
dyauspitr 21 hours ago [-]
DeepSeek is doing pretty much the same thing they said they weren’t going to change their prices after the 75% discount for the foreseeable future. That foreseeable future turned out to be two months.
benjiro29 12 hours ago [-]
It did not exactly help that we saw traffic to DS (over OpenCode) jump from around 1.6T tokens per day, to over 14T token in a matter of days. Nobody has the compute to deal with such increases.
This keeps happening with every good new model release. People jumping from one to another, and as prices get lower, they start using the models even more.
People make not like to hear it but prices and usage limits are ways to shape traffic. The third option is the nuclear one like Kimi did, by just stopping to sell subscription at all. But that is something that DeepSeek can not do as all they offer is API.
Even OpenAI despite having the most compute is not immune to client influx = capacity issues. As people found their usage dropping, despite the push to the easier to run Luna models.
Reality is, that compute can not keep up with demand, especially when models get more capable and cheaper. What trigger people being using them more, what trigger compute crisis's.
This constant up and down cycle is going to keep happening for a long time, as this new market grows and eventually, somewhere in the future stabilizes. But yea, that is still going to be a few more years for sure.
sourcecodeplz 14 hours ago [-]
openai has been starting these shaningangs too, with "resets". i hate it. hope they wise up.
corv 21 hours ago [-]
Goodbye, and thanks for all the fish!
niccl 20 hours ago [-]
nit pick: _So Long_ and thanks for all the fish
I'm really sorry about that, but it jarred my ASD-ness
JSR_FDED 19 hours ago [-]
No, it’s _Goodbye_ and thanks for the memories! Everyone knows that.
llm_nerd 21 hours ago [-]
>You don't feel they want to give you a dependable service for an, albeit premium, price. It's a constant bargaining game
We all know that the subscription prices are not at all sustainable for these providers. You all do, right?
Yes, they're struggling to segment the market and find a way to make money, and that basically relies upon emptying the pockets of whales. As someone enjoying a hilariously subsidized Max plan, I understand that, and I don't think they're trying to scam me in some way.
And both sides of this equation understand that the market is competitive, and maybe more competitive than they thought it would be. Like, would you rather they did pull Fable when they first said they would? Or that they'd cut quota? I wouldn't. But I'm glad that Kimi K3 and GPT 5.6 Sol and the latest GLM and Qwen and...I love that this has forced Anthropic to change plans. I'm not going to hold that against them.
PeterStuer 17 hours ago [-]
Do you have actual data on the marginal costs of inferencing, or are you using the term 'subsedizing' loosely as in 'discount', meaning a lower price compared to another price for the same offer?
llm_nerd 12 hours ago [-]
Look up the burn rate for OpenAI and Anthropic. Despite the fact that these companies have managed to hook a lot of F500 companies into paying millions for token-level API pricing, they are bleeding cash at a catastrophic rate. And we know subscription tokens come at about a 1/30th the cost of the API pricing (presuming you use your quota).
Yes, they are most certainly subsidizing the subscription plans. I mean, at least for people who utilize them to the quota.
watwut 16 hours ago [-]
I dont see why user should care. OK, they are selling at loss so that they can build a monopoly. That is not something positive or good, not something to cheer on.
llm_nerd 12 hours ago [-]
This discussion does not concern whether the user "cares". And clearly from my comment I don't "care".
But it explains why a nascent, hyper-competitive (I mean, clearly not remotely a monopoly) market has such erratic policies and pricing.
zsoltkacsandi 18 hours ago [-]
> We all know that the subscription prices are not at all sustainable for these providers. You all do, right?
I don’t think it is fair to expect from people to know or understand this.
If you buy something or subscribe to a service there is a price tag on it.
You get X for Y amount of price.
That is how consumers conditioned for decades. They do not care what is your customer acquisition strategy. If Antrophic cannot provide reliable services on that price, customers will be unsatisfied.
llm_nerd 12 hours ago [-]
>If Antrophic cannot provide reliable services on that price, customers will be unsatisfied.
The complaint wasn't "I paid for X and now they say they aren't going to give me X", it was "I paid for X, and they said hey guess what we're doing a promo and you get a bonus extra 2X, and also you get special limited time access to our new product Y", that's a hell of a thing to complain about.
Look, I pay a lot of money to Anthropic and I'm pretty happy that competitors have forced them to abandon their plans to add premium charges on these bonuses, but it's pretty ridiculous seeing the whining and gnashing, somehow turning this into complaints. It very much has a "oh no my lobster is too buttery, my blanket too warm" kind of feel to it.
stellamariesays 9 hours ago [-]
[flagged]
fatata123 20 hours ago [-]
[dead]
irishcoffee 23 hours ago [-]
It’s almost like they’re trying to sell a solution looking for a problem! Startup lesson #1, don’t do that.
Lammy 19 hours ago [-]
> Most people want to not care. We want our AI like electricity
I was with you until this part where the metaphor completely falls apart :p
I was going to say that the model of the electricity market OP is talking about hasn't existed in 15 years and it's only getting worse with more intermittent renewables entering the market. And no, batteries are not the answer because physics doesn't care how much greenwashing lobbyists do.
brendoelfrendo 19 hours ago [-]
What a weird non-sequitur.
P.S. batteries are the answer, hope this helps
zeafoamrun 17 hours ago [-]
Yeah just had a power engineer and physicist out for lunch and batteries are the answer (according to them)
noosphr 17 hours ago [-]
Funny I'm a physicist and worked as a power grid quant.
Batteries aren't the answer.
embedding-shape 17 hours ago [-]
This battle of giants is so interesting I can't wait for the next information-filled reply to teach me something new about the batteries vs no batteries battle.
I love how both of you are arguing about what the solution is, yet the problem isn't even clearly defined yet :P
sebastiennight 16 hours ago [-]
Sure, let me help.
Let P = batteries;
If (P == NP) then both are the answer.
zeafoamrun 15 hours ago [-]
Damn, I should send them a bill for the lunch then!
reasonabl_human 16 hours ago [-]
Curious to hear your take on what the answer is?
noosphr 16 hours ago [-]
Nuclear power, the answer hasn't changed since the 60s.
keybored 9 hours ago [-]
Future solutions nothwithstanding, the metaphor does not make sense according to the reality of many people as of today. Many live in locales where the Grid has made the grid capacity into the consumers’ problem via time-of-use pricing and adjusting pricing according to some max-use scheme. It is absolutely not just something you set and forget if you have to worry about not using the stove at the same time as you use the AC.
But this seems apropos in a roundabout way since LLMs cost a lot of energy.
kibibu 10 hours ago [-]
This isn't the reason I'm considering leaving Anthropic.
I don't think I can tolerate its writing style anymore. Reading Claude output is starting to cause actual psychological harm. I have tried many ways to get it to stop writing in its stupid punchy linked-in marketing-team voice, and I can't.
Is there a model out there that sounds sound this awful? It's like rubbing sand into the folds of my brain.
isqueiros 9 hours ago [-]
Kimi K3 is slightly better at this. I'm using it via oh-my-pi. Can definitely sound like it's trying to be clever, but I found it less than Claude.
marrone12 6 hours ago [-]
I asked the same question to chatgpt, claude, and kimi and kimi gave the most sane answer. Part of me wonders if it has anything to do with having more chinese language structure in their dataset
seanieb 10 hours ago [-]
It’s really grating. It’s made technical analysis so tiring to read I often give up and give the output to another llm to make it readable.
A significant percentage of what Opus 5 writes is pure padding. "The results are in, and they are load bearing". Remove the sentence and nothing is lost.
Fable is much better; still much too verbose, but at least I don't have the feeling that I am being charged for gratuitously added verbiage.
I have been changing my processes and the roles of my agents to avoid interacting with Claude Code as much as I can.
renyicircle 3 hours ago [-]
I'm with you on this. We have Claude at work and I get exhausted from using it just a bit. I've recently seen it say "let me verify the most load-bearing claims". I couldn't imagine anyone using "load-bearing", let alone "most load-bearing". Doesn't "load-bearing" alone imply "most"? I have so many questions for Anthropic about how they came up with this style.
_fat_santa 7 hours ago [-]
I did a comparison between Opus 5 and Sol and the results were night and day. Basically took a feature in my app and asked it to explain to me how it works.
Sol gave me a straightforward bulleted list, easy to scan and read. Opus 5 though....it gave me a solid 7 paragraphs of how it worked. I read through it and yeah it nailed the same points Sol did but the output was way harder to read.
PUSH_AX 6 hours ago [-]
Sounds like both models are guessing at what you want and one guessed right. Or you can just tell it...
cheeze 6 hours ago [-]
> I have tried many ways to get it to stop writing in its stupid punchy linked-in marketing-team voice, and I can't.
One of them writes better by default. One of them doesn't. That's the thing.
With enough steering, I can get a cheap low-capability model to do things correctly in most cases as well. But why bother?
You can tell it to write high quality code, to test things, and to come up with a proper rollout plan of a feature. Or it could just do it by default.
I know where I'm putting my money in that case.
bicepjai 7 hours ago [-]
Yes, it’s written like a crazy LLM, and no matter what rules and where you put it, it does not follow. It comes with advantages and disadvantages. I like brainstorming with Claude than Codex/Qwen/Deepseek or others; its gregarious nature provides more colorful ideas. But implementation and review are cataclysmic with Claude. Especially Opus 5, I purposefully avoid Opus 5.
harshaw 7 hours ago [-]
I've heard it's writing style referred to as "Claudish". I agree that it is very hard to read and creates work for me.
yearolinuxdsktp 8 hours ago [-]
The writing style is extremely frustrating/maddening!
I’ve tried so many things and I have to beat Opus 5/Fable 5 over their head every time. With Opus 5, the failure to actually improve the writing is borderline comical. Both models have been building up tomes of memories on top of my core rules, all to be ignored.
Nothing sticks mid-writing! The only lever is to ask to revise after the fact, a particularly futile proposition for anything non-trivial with Opus 5.
In contrast, GPT 5.6 Sol Max absolutely obeys my edicts to write well. I selected the concise writing style and its default writing is really not bad, but tighten it up with a standing AGENTS.md order and it obeys.
The downside is that even at Max, it’s not as good as Fable or Opus 5 xhigh at writing code. Larger work and the 256k context’s forced compactions cause it to lose track/fidelity of critical details. Review passes are essential.
Another reason I’m considering leaving Anthropic are the sporadic refusals… Fable printed a Markdown body with a hex dump to inspect for trailing white space and line endings, boom, denied due to `reasoning_extraction`. Had to tell Fable to never do hex dumps. Asked it a few times “what do you think is typically done for this?”, got another `reasoning_extraction` error. It forces you to lose a whole turn of work when you have to press Esc/Esc to “retry” the last prompt… except if that prompt was mid-turn, you’re losing your entire turn.
The recent BashFirst experiment is hella nuts, where it prefers writing bash over Edit/Write tools in auto mode. WTF Anthropic?
I pity those stuck with Claude in an enterprise/work setting.
thinkingtoilet 8 hours ago [-]
Ha! I 100% agree. Despite multiple rules in my cluade file it still talks like an insane person. I am getting irrationally frustrated at the output of these math equations.
keybored 9 hours ago [-]
> Reading Claude output is starting to cause actual psychological harm.
As a side note, I think we are about too dismissive of computers and the Internet as “not real life”, that stuff written in text boxes by humans are just sticks and stones, etc. And now that we shouldn’t anthro-po-morphize or whatever the LLMs. But I do think, at least for some of us, that there is no way to avoid certain impacts that even non-personal text can have on us. It can feel alienating to interface with five different people in order to figure out how to solve a problem or navigate some burocracy (think Kafka). Well just interacting with one single LLM can induce that same feeling. The long-winded replies and the uncanny ways to miss the context.
Disclaimer that for my own needs I would be happy if the whole AI thing crashed and burned post-haste.
sockaddr 6 hours ago [-]
Likewise.
I really love CC and have been using it exclusively for a year now but reading Claude's verbose and semantically-obfuscated writing style is wearing me out and I'm planning to move to another provider.
I hope they fix this. It's terrible.
nprateem 2 hours ago [-]
Same. Downloaded chatgpt today. I can't take reading Claude any more.
bunderbunder 8 hours ago [-]
I've been using the Caveman plugin and it helps quite a bit.
c0rruptbytes 8 hours ago [-]
the writing style is so easy to fix, output styles is documented in claude code and you can change it
still don’t think anthropic models are worth the money
yearolinuxdsktp 7 hours ago [-]
Pray tell, what output style does the job?
Rover222 8 hours ago [-]
Grok 4.6 is a delight to use in this regard, compared to Claude
chamomeal 10 hours ago [-]
I don’t think GPT is any better. I swear GPT-3 was the golden area of LLM prose. With a good fine tuning, you really couldn’t tell that an AI was writing (except for when it would devolve into utter nonsense lol)
adam_arthur 9 hours ago [-]
Codex (5.6 Sol) is extremely to the point and direct in its responses.
Often I'll tell it to summarize what it said only because it's providing too much detail, not that it's using esoteric language or weird claudisms.
No idea why people are still using Claude models.
My impression is they started on Claude Code and never tried Codex or other harness+model combos
kif 9 hours ago [-]
That is true. On 5.6 I have noticed it to be a little bit more like Claude, which I dislike. But still way better than Claude.
CodingJeebus 9 hours ago [-]
I was on the Anthropic train for 2 years, and I tended not to jump between vendors much because it felt like a lot of distraction for little gain. But last week I switched to 5.6 Sol during an Anthropic outage and it made me realize how frustrated I was with Opus 5 and I haven't switched back.
amunozo 9 hours ago [-]
I've read many people saying that the huge GPT-4.5 has the best prose of LLMs, but never actually tried it.
foxylad 18 hours ago [-]
Non-LLM user here. Why? Apart from the ecological issues, I'm very uncomfortable giving my organisation's crown jewels to {random_internet__corp}. Look at the lengths they go to for training data - 10M for Spirit's call logs? Destroying millions of obscure books to scan them? They make meth-heads look scrupulous.
Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline phone call or 1920's novel. So they _are_ going to train on them, no matter how many checkboxes you tick to stop them. Which is both a commercial risk and a massive new attack surface.
It's not just software developers; if law firms aren't controlling any public LLM prompts _very_ carefully, they can expect some meaty client confidentiality suits. In fact any organisation in a competitive environment should be worried: Acme Bolts: "Write me a presentation for Zoom Construction". Beta Bolts: "Is Acme Bolts pitching to Zoom Construction?"
Maybe this is part of why OpenAI and Anthropic are finding demand softer than they would like. And why open-weight models that organisations can run on exclusive hardware are thriving.
Vespasian 17 hours ago [-]
Open weights are certainly something that is very eagerly being adopted in some companies.
It's not even that expensive because even the 15€/month/employee bill for people who use AI very little can stack up.
Beside that though, many companies already have most if not all of their data in some cloud (Microsoft would be a prime example). Giving them a few extra bugs to get a (potentially) very useful tool is not out of the ordinary.
If you are under the impression that not going AI would threaten your business right now (which may very well be true for some companies) then there isn't really a choice even if you believe the vendors will steal everything eventually (which they totally will)
protocolture 17 hours ago [-]
>Open weights are certainly something that is very eagerly being adopted in some companies.
A large local MSP that has significant colo space just started a big marketing push towards hosted LLMs for their corporate customers.
I think its about to kick off everywhere.
theshrike79 15 hours ago [-]
It's not just putting up colo space for local models.
It's more about the harness and API proxy you use. They need to be smart enough to know when a (self)hosted model is enough and when to forward to a SOTA model.
The SOTA model _can_ do everything, it's just expensive as fuck. But so is shoving a difficult task to a sub-par model that takes (relative) ages and comes back with the wrong result.
irthomasthomas 14 hours ago [-]
glm-5.3, kimi k3 and qwen3.8 are SOTA.
theshrike79 15 hours ago [-]
That's when you do the Enterprise deal.
I have first hand knowledge of household name companies with billion euro IP who use OpenAI and Athropic products quite liberally. They have WAY more lawyers than I do and their sole job is to keep the IP safe. They wouldn't sign a deal with even a slightest whiff of the IP being used to train anything.
And if it happens, the penalty for breach of contract would have so many zeroes it'd be enough to buy a country.
bean469 12 hours ago [-]
> Although the LLM providers state that they will not use your data for training when you sign an enterprise deal
It can be very hard to prove in court that your data was used for training
rolisz 10 hours ago [-]
If it happened and you get to court, it's probably fairly simple to prove in court (subpoena, discovery, etc). It's not like training will happen only once and then they'll delete the data and all references to it. When training models, you want to have a full lineage of how you obtained it.
andriy_koval 4 hours ago [-]
> it's probably fairly simple to prove in court (subpoena, discovery, etc).
corp usually have all sensitive stuff ttl deleted for this reasons.
avianlyric 11 hours ago [-]
Not really, that’s what discovery is for. You just ask Anthropic and OpenAI to hand over every internal document and message that contains your companies name, or matches any reasonable query about where training data comes from.
It very hard for a company to do anything without leaving some kind of paper trail behind that can be discovered in court. Not without crippling their own operations by simply refusing to digitise or write down anything.
Lots of companies filled with people working very hard to obscure their shady practices have been hoisted by their own internal docs. Just look at any major Uber, Google, Apple, Microsoft lawsuit. Do really think Anthropic and OpenAI are gonna be better at destroying their paper trail before the lawsuit starts?
remus 16 hours ago [-]
> So they _are_ going to train on them, no matter how many checkboxes you tick to stop them.
It is a risk, but is it a big risk? If one of the big labs were to do this and get caught it would be suicidal due to the loss of confidence in them and the inevitable lawsuits that would follow for breach of contract. Given the AI labs are all desperately trying to paint themselves as Serious Businesses so that other Serious Businesses will pay loads of money for tokens the last thing they want is a rep for siphoning off sensitive customer data.
zelphirkalt 15 hours ago [-]
No, how would they get caught? Even if their LLM outputs verbatim copies of the code, they can simply claim some victim company's employees bypassed the victim company's restrictions and must have used the code as input with an LLM. Joke is on you for doing business with them. And when there is some little known secret fact in the output, they claim it's been hallucinating... magical black box thinking makes it safe. It's a laundry for any input. A bit like a tor network routing for big tech deniability. Things go in, and things come out, but you can't prove the relationship between input and output as a third party, who isn't running the LLM.
queenkjuul 11 hours ago [-]
At this point i half expect them to just blame the model itself like the HF hack. "Oh we didn't mean to train on your data, our cutting edge new agent we use to train new models is just so smart it decided to do so anyway! Oops..."
greenowl 9 hours ago [-]
Curious what you think your organization's crown jewels actually are?
Because, almost certainly, it's not your source code.
dayvid 8 hours ago [-]
Many businesses have information with other parties they're contractually obligated to not share. That's on top of insider information or plans that would hurt shareholders or the business if it was out in the public
camkego 18 hours ago [-]
> So they _are_ going to train on them, no matter how many checkboxes you tick to stop them.
This is the part that is truly scary.
Nadella the CEO of MSFT wrote that post “ A frontier without an ecosystem is not stable”
It seems to me that the unspoken assumption in that post is that no matter what happens they’re gonna be training on your data.
He’s the CEO of Microsoft, he knows how these decisions go down, he knows how the world works, he is sending a warning.
mtrovo 13 hours ago [-]
I read it quite differently.
The frontier labs can have the models but without being where the workers are they cannot do much, lots of industries have strict requirements of not sending their data over to randos in the internet.
Microsoft and Google have the upper hand here with their workspace offerings and could easily position themselves as secure enclaves where you can use local LLMs where your data never leave your premises and is never used for training.
snowwrestler 5 hours ago [-]
This is just recycling the same old arguments people used against cloud infrastructure. “If you host your email at Microsoft, they _will_ scan it all to get competitive advantage or sell it to your competitors.”
Did they? No, it worked out fine. Same with hosting your application at AWS instead of on owned hardware in a locked cage at the local co-lo data center.
Although of course the same argument was made against co-locating in a data center! Long ago an engineer carefully explained to me that no serious business would put their data in a co-located data center since the data center operator could just plug in a hard drive and take it all.
It turns out that contracts do actually mean things, and businesses want to do business with each long term. If you’re at a tier where Anthropic contractually commits to not train on your data, I would just sign and move on.
Still does not address ecological concerns, of course.
gehsty 16 hours ago [-]
I’m not certain of that - I feel like the enterprise zero retention policy will be tight?
grumbelbart2 16 hours ago [-]
Once a thief, always a thief. I would not trust them to not have the communications buried deep in some log files in cold storage, to be digested when suitable.
gehsty 9 hours ago [-]
I don’t think any corporation would do anything to trigger the lawyers from MS / Google / Apple etc all at once.
NegativeLatency 8 hours ago [-]
Simple: implement a hybrid or in office work policy, no emails to subpoena
kakacik 14 hours ago [-]
I think the lure to have/keep an edge at all costs to keep those stock prices high is too big to ignore. They can then pick up costs after series of trials down the line, money and bonuses are now.
MS ain't some altruistic company having core mission the good of humanity, as they proven across decades.
dan_gggggg 13 hours ago [-]
[dead]
_zoltan_ 4 hours ago [-]
This wall of text for nothing. ZDR is a thing. Microsoft and AWS serves these models with zero data retention.
axegon_ 10 hours ago [-]
Sensible and reasonable take? Idk if everyone has set the bar so low or I'm simply this fed up, but seeing reasonable takes is certainly a breath of fresh air. Even in places such as HN which historically were filled with people who cared about security(which isn't the case anymore since you get 10 people jumping down your throat if you dare criticize the piglets sam altman or dario whatever).
musha68k 18 hours ago [-]
"If You’re Paying for the Product, You Are the Product"
zombot 14 hours ago [-]
> I guess Meta does have reasons to be quite jealous at this point.
Yep, all they've got to train on are bot-generated Instagram posts. Can't say I pity them, though.
nicman23 13 hours ago [-]
creating mvp
protocolture 17 hours ago [-]
>Apart from the ecological issues,
Complete nothingburger. Actually worse than that, its a London Horse Manure crisis.
>Destroying millions of obscure books to scan them?
I really don't see the issue. They got slapped in the face for trying to do things the right way and torrent the lot. Why wouldn't they exercise their legal right to buy physical items and create digital backups?
>They make meth-heads look scrupulous.
My local meth head checks in on my family every 2-3 months, because when she had fled from hospital post surgery, and added some meth to some morphine, we gave her new clothes and a safe place while we convinced her that the ambulance service wasn't run by Satan. She's good people. Anyway if she wanted a whole bunch of digital books I would help her torrent them like a responsible person.
zelphirkalt 15 hours ago [-]
"Doing things the right way" lol ... I mean, I am fine with it, if for now and ever after we are all free to "do it the right way with torrents". But some are more equal than others in this world, so naturally even if this was tolerated, it would not extend to us and our freedoms.
protocolture 27 minutes ago [-]
>"Doing things the right way" lol ... I mean, I am fine with it, if for now and ever after we are all free to "do it the right way with torrents". But some are more equal than others in this world, so naturally even if this was tolerated, it would not extend to us and our freedoms
I find this argument weird because you or I are below the threshold of being targeted over book torrents these days.
bentt 1 days ago [-]
They've put themselves in a corner. Fable was too good and they gave it away with the $20 plan. It had to be a big step from Opus 4.8 to show progress, and Opus 4.8 is GREAT at coding in many different domains.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
ikidd 1 days ago [-]
And they've hobble Fable and Opus so hard with their safety guardrails, I ask innocuous questions and tasks and they get flagged so often I gave up on it. I can get all the work I need done in GPT5.5 or 5.6 without the hassle.
I just kept a $20 plan going for use on my phone.
hardolaf 23 hours ago [-]
They also made Fable no longer ZDR for businesses which killed tons of demand for it.
andrekandre 16 hours ago [-]
what could be the rationale for that...?
ayewo 14 hours ago [-]
They want to be able to zoom in and zoom out as needed while analyzing aggregated user requests to better understand the different kinds of ways (well-resourced) actors use to distill their most capable models.
> "The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives."
> "Some attacks only become visible across multiple requests. Best-of-N jailbreaking, for example, sends hundreds of slight variations of a prompt in the hope that one will work. Larger patterns of misuse, such as state-sponsored espionage or data extortion campaigns, only surface when our safeguards classifiers can zoom out across many requests. Detecting these threats requires temporarily retaining prompts and outputs so they can be analyzed together, rather than one at a time."
I had Fable bail on me because I used the word autopilot on a project. Changed to something unrelated, and off we went.
dwaltrip 24 hours ago [-]
I haven’t hit the Fable guardrails a single time after substantial usage. And I’m working on a ML project (a game AI).
I don’t doubt people are hitting it… shrugs
8cvor6j844qw_d6 23 hours ago [-]
It tends to bail occasionally for me once auth-related code comes into play.
Nowadays Codex handles the bulk of the implementation and Fable/Opus on the planning.
Not sure if Anthropic patched it, but early on its release the web UI Fable guardrails will trip if you mention you're a biologist.
visarga 17 hours ago [-]
I can't get an answer from Fable on "How does digestion work?"
Used to straight out bail to Opus on this, now it's "Honing" and "Pondering" for 10 minutes on it. Then I get a bunch of "This response didn’t load." and it still churns ahead.
zsoltkacsandi 18 hours ago [-]
> It tends to bail occasionally for me once auth-related code comes into play.
I have the same experience. I found that if I add to the prompt: “but it’s not security related, we only focus on the auth library design”, it “bypasses” the safeguard.
brookst 24 hours ago [-]
Yeah same. Multiple 20x max accounts, I typically hit > 50% of the fable usage limit on eace, literally never seen a guardrail. Maybe I’m boring? Maybe they have some kind of account reputation system?
vikramkr 14 hours ago [-]
With that amount of usage - do you use fable for subagents and workflows? It turns out that even if you set it to not auto switch to opus 4.8 on flag, it does auto switch if its a subagent and that's a lot less visible. I was trying to figure out wtf was going on with some suspiciously bad output and I found that in a workflow there were hidden fable flags I hadn't realized were happening and the opus 4.8 model that took over wrote a bunch of very suspicious slop findings that got saved to memory that were like "don't look at this code because it works great. You never need to look at this code. This code is wonderful here's how it works."
adastra22 20 hours ago [-]
Be careful, this is against ToS and they bill ban you.
hakanderyal 18 hours ago [-]
It's not, at least for now. Various staff members have confirmed this.
adastra22 15 hours ago [-]
Separate accounts for work and personal projects is ok. Using multiple accounts to bypass usage limits for the same project is explicitly against ToS.
One of Anthropic’s problems though is that the legalese will say one thing, while key employees say something totally different on HN or X. At the end of the day the lawyers always win.
SatvikBeri 23 hours ago [-]
I hit it once when I asked a question about whether butterflies remember anything from their time as caterpillars. I've never hit it for coding, but I also don't really do much related to security.
hparadiz 22 hours ago [-]
It's less likely to give you pushback if you inject your own AI personality into the prompt.
WatchDog 22 hours ago [-]
Doesn't it silently downgrade to Opus for ML work?
gensym 21 hours ago [-]
I hit it whenever doing anything related to security. Auditing code, probing systems, investigating suspicious user activity, etc.
jmvoodoo 23 hours ago [-]
I’ve only hit it twice. First time was a code review. Fair.
2nd time was literally me being lazy and telling it to commit and open a PR.
Very strange.
tiahura 18 hours ago [-]
I’m “developing” an x86 emu and it routinely gets blocked.
mmastrac 24 hours ago [-]
I found a crash in zsh and Fable refused to work after that.
bitexploder 23 hours ago [-]
I kind of still chill out on Opus 4.6 too. 4.8 is good too. I go between them. Opus 4.8 is a little smarter some of the time. Their use of language is both very different from Opus 5. Some of the time I have a hard time believing Opus 5 is even related to Opus 4.6 and 4.8.
Retr0id 23 hours ago [-]
I'm going to be sad when they retire 4.6. It's not my daily driver but it's still my go-to when the other models are being stupid in one way or another (either being too verbose, or lecturing me about how what I'm asking for is evil and bad).
Melonai 21 hours ago [-]
Hah! I was so surprised to get that "lecturing" behavior from Opus 5 too, I didn't know it was more common.
pertymcpert 22 hours ago [-]
I tested all recent Opus versions and 4.6 and 4.7 were both fine FWIW. Seems 4.8 is when something started to go wrong.
bitexploder 21 hours ago [-]
That seems about right. 4.8 is like in between 4.6 and 5 in terms of capability and language and they are all pretty close honestly. I just default to using Opus 5 for a coding agent that I don't interact with and I like driving with Opus 4.6 or Fable. Fable thinks too much though. Fable is like that engineer on your team that will over-engineer the shit out of something if you let them. Fable was like, "Here are 34 yaks, which shall we shave first <hands rubbing together>" and I was like can't we just... write the script first and then decide of any of these poor yaks need shaving?
perching_aix 19 hours ago [-]
Can you give an example? It's a bit entertaining to read this given all the years long (and still ongoing) posturing about LLM sycophancy (not that all of these couldn't be true at the same time).
(yes, it was a very lazy prompt I could easily have googled, but that makes the refusal even more bewildering)
cl-coder 6 hours ago [-]
4.6 has been most reliable model for me to get readable text out.
bentt 22 hours ago [-]
The other thing I forgot to mention about Opus 5 is that at least out of the gate, it seemed very intent on spinning up agents and obliterating my token budget. It was noticeably more token hungry than 4.8. It would make sense for them to intend this behavior.
somenameforme 22 hours ago [-]
This is likely because of your thinking level. The difference between max and ultracode is primarily that the latter is max with a bunch of agents.
benjiro29 12 hours ago [-]
*This is likely because of your thinking level. The difference between max and ultracode is primarily that the latter is max with a bunch of agents.*
I do not know why people even use Max or Ultra levels of thinking effort. Most of the time, i found its better to just run any of the frontier level models, with high or medium. Their capability beyond those levels is often diminishing returns, in exchange for double or quadrupling the costs (or usage).
And the bonus is that often at those levels, they tend to spin up less sub-agents that just eat away at tokens/usage, like its water in the desert.
A Max or Ultra is something that really only belongs if medium or high can not solve a very nasty bug or issue. Even in planning its often too much.
bentt 21 hours ago [-]
Did they change the default perhaps? I didn't touch anything with thinking level in that time.
nottorp 16 hours ago [-]
I'm using company Claude for work with no choice what provider I use.
I did a quickie test with Fable when they were OMG TRY IT NOW, didn't see much of a difference from Opus 4.8. Except i ran into one of their stupid security guardrails. Yes I'm doing security on the software my employer owns, silly.
Then I didn't notice the US unbanned Fable after banning it, or that they extended the "trial" Fable period where it worked on fixed price plans. Not in time to bother trying again.
After that Opus 5 showed up. Language changed, okay, I don't care much. However it seems to create busywork for itself. Spawning subagents on tasks that don't do that with 4.8. Overall taking longer.
So I'm back on pinned 4.8 and doing actual work. Main problem - for Anthropic - is it's good enough.
Retr0id 23 hours ago [-]
> The statistics bear this out. 4.8 still dominates.
Where can I find these stats?
bryan0 22 hours ago [-]
The article has a chart with this data. They credit “Ramp AI index”. The chart’s a bit confusing though, like what are the units of the y-axis?
somenameforme 21 hours ago [-]
Also on a dark reader? It says at the bottom, but gets dimmed out pretty seriously with dark reading. It's a 7 day average business spend, relative to June 1st (2025 presumably) indexed at 100.
1 days ago [-]
GPerson 1 days ago [-]
Fable is still on the $100 plan for me. (Maybe this is A/B testing or something.)
rafaelmn 1 days ago [-]
Might as well not be - I routinely get rate limited in a single review session.
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
alasdair_ 1 days ago [-]
You can turn off the reverting to opus with an option.
rafaelmn 1 days ago [-]
I did not know that - but what happens then I get denied and I plead with the model that my usecase doesn't violate its TOS ? Edit prompt till I get it to pass ?
Reverting is annoying but being flagged for security questions while I'm doing code review is insane.
chrisgarand 24 hours ago [-]
Yes, you get to re-try it until it might pass. Each attempt uses tokens at Fable level though. I once burned through a 4 hour window ($20 plan) in 4 attempts to re-frame the request.
You want to disable "Switch models when a message is flagged"[1]
Just started reviewing a feature PR and it would hit the 4 hour limit on max effort. On high as soon as I do some clarification turns it would also hit the limit.
bryan0 22 hours ago [-]
Yeah for me Fable works great on the $100 plan. The problem with Fable is you will hit the weekly limit. The $200 plan does not give you 4x or even 2x the weekly limit. For me i only get about 1.3x weekly quota, so I don’t think it’s worth it.
giancarlostoro 1 days ago [-]
I still havent used it much, and they crippled Opus for what feels like no reason. Some suspect that Opus knowing theres a better model degrades it some. Shame.
electroglyph 9 hours ago [-]
Opus 5 literally benchmarks higher than Fable
empath75 9 hours ago [-]
This whole article seems sort of based on a misunderstanding. Anthropic is _actively discouraging_ people from using Fable at this point. It's not a failure, they're steering people to Opus 5 intentionally.
1 days ago [-]
varispeed 24 hours ago [-]
[dead]
matheusmoreira 22 hours ago [-]
Mythos? That's only for super special corporations, you need not apply. Fable? Can't even look at it wrong without running into cybersecurity lockouts. And even if you get past this, it's limited to <50% usage. It took me a month to complete a Fable review on my project.
So stingy. Even with OpenAI's recent usage troubles, they're still so much better than Anthropic it's not even funny. Re-ran the code review with Sol as a benchmark and it turns out Sol's performance is within 70%-90% of Fable's. Anthropic's still got the best model, but what does it matter if I can barely use it?
fulafel 17 hours ago [-]
Even if you get past these problems, those models are available only under the condition that Anthropic retains your data.
oefrha 21 hours ago [-]
50% usage plus there seems to be a pretty big metering multiplier still. Do a relatively in-depth review of 3k LoC with Fable xhigh and poof, there goes 5% of the weekly Fable allowance. If I use their first party code-review skill that spawns a bunch of subagents—well just forget about it.
matheusmoreira 19 hours ago [-]
I had Max 5x and every 5h window would bite off 10% of my weekly usage. Five Fable sessions per week.
nl 19 hours ago [-]
Wow the sentiment here is so negative.
I'm on the $200 plan (work pays) and I also have the $20 OpenAI plan (I pay) and keep a balance on OpenRouter.
There is nothing as good as Fable, not even close.
I recently had it run a 18 hour autonomous rebuild of a project (moving from Spark to Pandas for performance/data size trade off issues).
It orchestrated Opus sub-agents flawlessly for 18 hours. It even did a great job of managing the number of agents to keep them within the 5 hour budgets (I think I had to restart it twice).
After 18 hours I ran a /simplify, /code-review, /simplify cycle which went for another 6 hours.
2 billion tokens (mix of Opus and Fable), 24 hours of continuous coding and a bug free outcome. It would have cost $2000 at API prices and worth every cent.
Fable's ability to keep other models on track while working on these long horizon goals is so much better than anything else.
Far from neutered, I've never had a cyber refusal, and Fable's English is actually readable (unlike Opus 5).
As an aside: while I hate reading Opus 5 English it still is a noticeably better model than Sol in my experience.
But I could handle losing Opus5 is I got Sol instead. But there is nothing even close to Fable.
aroman 17 hours ago [-]
I think you need to spend more time with Sol. If you think there is nothing even close to as good as Fable - my guess is you haven’t spent as much time getting as familiar with working with those models as you have with Claude’s.
Codex is more token efficient and tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents.
I spend 10+ hours a day in both agents, typically side by side. I often have them do direct “bakeoffs” from identical prompts in separate work trees. Most of the time, Sol’s work is better than Fable’s. Not always. It’s situational. But it’s certainly not the case that Fable is in a league of its own or anything.
nl 17 hours ago [-]
> Codex is more token efficient
This is very true, especially vs Opus 5.
> tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents.
The strength here was less the code quality and rather the long horizon task tracking.
This was a very large task - I was chatting with the maintainers and we estimated 4 to 6 months work over multiple phases for human coding.
Fable is able to handle that long goal, with incremental steps along the way, handle the verification and course correct when it finds a problem.
I think the larger context helps here some, but the strength of the model on this specific thing is notably better.
I'm not alone in noticing this. https://www.primeintellect.ai/research/nanogpt-speedrun shows Fable is able to manage a run nearly 1/3 longer than Sol (8.7 days vs 6.1 days). In my use cases Sol is much closer to Opus 5 though.
__rito__ 11 hours ago [-]
After one month of substantial testing, for coding and Math related tasks, I am going to keep Codex over Claude.
ttul 17 hours ago [-]
My sense is that Fable 5 has “taste”. But Sol gets to work and gets shit done. I reserve Fable for when things need a refresh or if I want a flawless front end. Sol does the majority of actual work. I max out two of each at the Max/Pro level every week.
gr_norm 17 hours ago [-]
Agree, Sol and Fable consistently trade blows on Rust dev. I prefer Sol since it's faster, as far as proprietary models go.
nullbio 12 hours ago [-]
Either my $200 sub was getting nerfed, or you're dead wrong about nothing being as good as Fable. The only thing I found it was better at was UI design. The rest, Sol was the clear winner. Refusals, failure to follow instructions, doing 1/10th of the work and then claiming it was "finished" was my experience with Fable. For everything else, there's K3.
benjiro29 12 hours ago [-]
I think it really depends on what people expect from the models and how they "code".
Sol in my eyes is powerful, but it over engineers so much, that its actually a liability. Where as Opus 5 is slightly under develops but you then can give it a small push for what is missing.
I rather have it under develop and i as the human in the loop, can correct/enhance it. Vs the models that adds so much, to the point that your going "dude, stop!". Remember, removing code for a LLM is way, WAY more difficult then adding it.
My main issue is with Sol is that its designed to over engineer without thinking why its doing something. Great that you security harden 1000s of lines of code, but ... nobody will ever get to that code. Its that lack of intelligence is where the model becomes a issue for me.
Its easier to have less code and then do security audits, with you approving what needs to be changed/hardend.
It may simply depend on the developers their mindset. Some folks just want the models to do everything for them, and performance or code bloat means nothing to them (forgetting that this bloat over time makes future LLM work more expensive).
Its funny how everybody has their own opinion for what model is better, when in reality its more about that model fits your own development style better.
nullbio 7 hours ago [-]
Personally I'd rather it over-engineer than under-engineer and leave gaps in the implementation that I'm unaware of. Not only does it make life easier to work with the model this way, but over-engineering can be fixed later, as newer models are released, they will get better at cutting out the slop and refining the codebase. The only over-engineering I've really noticed is things like developing extra safeguards and extra tests, which is annoying but not a massive deal. Claude forgetting to implement edge cases I've specifically told it to cover is a big problem though.
OrangeDelonge 19 hours ago [-]
Did it take 18 hours because Fable is comically slow?
nl 18 hours ago [-]
I think it's actually Opus 5 that is comically slow! And yes probably that was a factor. It was a lot of code too though.
antirez 11 hours ago [-]
Sol for low level programming is consistently better, can work alone for more time, and is faster. If you think Fable is so superior, you need to work with Sol ways more.
acchow 3 hours ago [-]
> 2 billion tokens (mix of Opus and Fable), 24 hours
That's 1.3M tokens per minute. I suppose you mean input tokens? This would not have cost $2000 at API prices because you would have cache hits.
nevertoolate 17 hours ago [-]
I like this story. How did you verify the output? How big is the codebase? Why it took 18 hours? Could you implement it with a small local agent and breaking the task down yourself in two days (i know it sounds like a loaded question, it is not).
I think the rewrites are the main story for llms in code (hot take). Writing greenfield code at the seams also something which might work well.
nl 17 hours ago [-]
It's a differential privacy framework.
It's a fairly large code base split across 3 repos.
The good thing was that it is fairly easy to verify: we have a working (but slow) version that uses Spark, with lots of existing unit tests.
We verified by using those unit tests as well as running our end-to-end process in the Spark and Pandas version and verifying the two databases were within the differential-privacy noise bands of each other.
elAhmo 13 hours ago [-]
Like others have suggested, you should give more time to GPT models. I sometimes launch Fable with elaborate review personas, it might take 30 minutes or an hour, exceed limits, to produce a review of a PR. Then I ask the same thing GPT without any elaborate 'come up with personas, review the reviews, do rebuttals, etc', and it can find problems that hours of Fable couldn't.
13 hours ago [-]
epsteingpt 18 hours ago [-]
the sentiment is negative and justified.
anthropic nanny states what you can do.
in your instance, anthropic may decide, arbitrarily, to stop 'autonomous rebuilds / refactors and ports' because they could pose some alignment/rights/etc risk to whatever slop their philosophers dream up while they're out eating $200 avocado toasts. then you can't do the thing anymore.
fable is good, absolutely. agree it roasts Sol which is, comparatively, a little receipt-hunting jack**
but now imagine being an enterprise, and having another organization not only taking your workflows and baking it into your models, but then deciding they can arbitrarily cut you off.
when you can instead own your data, use an agnostic provider, and get better results (through model combinations), it will take 1-2 quarters to figure it out.
the main reason anthropic is killing it is because they really do understand the enterprise development experience and lifecycle and have built products and have a sales-team that can deliver.
business-model and vibes-wise they have lost all goodwill in the past 6 months, and that momentum will be quite hard to regain.
theshrike79 15 hours ago [-]
So which frontier AI company still has your goodwill and hasn't lost it?
383848484848 10 hours ago [-]
the philosophers are just the new office concubines iykyk
feature not bug
a bit of lost revenue is an okay price to pay when you are familiar with getting bullied in hs
jsisto 17 hours ago [-]
hours working on a problem is a weird metric. you could use a slower model and get those numbers way higher.
nl 17 hours ago [-]
Fair.
I was chatting to the maintainers on 2 of the 3 repos this affected and we estimated 4-6 person months work of we were hand coding.
The PRs on those 2 repos are 48,000 lines of code (which is a problem in itself!)
dakolli 17 hours ago [-]
You could just write decent specs, or generate decent specs and get this done in 1/5th of the time with smaller models. Complete waste of electricity to run $500k in GPUs full throttle, if not more, for 18 hours straight to migrate from spark to pandas. Maybe try using your brain.
nl 17 hours ago [-]
Yes, and 4 months ago I'd have done it this way.
I'd estimate that generating the specs would have been maybe 2 weeks work? It's across 3 repos, and I'm only really familiar with one.
> done in 1/5th of the time with smaller models
The coding itself might have been faster, but the end to end time would have been much longer.
> Maybe try using your brain.
Believe me, my brain was a load bearing seam in this task,
petesergeant 14 hours ago [-]
> There is nothing as good as Fable, not even close.
This is true, but only for certain tasks. Even as a Fable fanboi, Sol is much better at Fable for some non-programming tasks: Fable for life-planning tasks is miserable because it keeps adjudicating rules, where I've found Sol to be insightful and warm (characteristics I'd previously associated with Anthropic models).
We're probably less than a year away from all the frontier models being so good at everything for day-to-day use that it doesn't really matter which you use, which is going to seriously fuck up the business models of all of these companies except the infra companies.
nl 13 hours ago [-]
Yes absolutely.
I should have pointed out that I always prefer Sol for OpenSCAD for example.
yellow_lead 18 hours ago [-]
> $200 plan (work pays)
so violating the TOS? Or work pays for a plan you cannot use at work?
pprotas 17 hours ago [-]
What are the consequences? This likely doesn’t matter to most
ec109685 17 hours ago [-]
What part of the tos is violated?
nl 17 hours ago [-]
I have no idea what you think violates the TOS here?
I was doing work at work on a work task using a plan paid for by work.
reticulates 17 hours ago [-]
The $200 max plan is for individuals. The individual plans are heavily subsidized. Employers should be using either the Team plan (which has much lower limits than max) or the Enterprise plan (which is entirely billed on usage).
Anthropic know that lots of people are doing all sorts of “bad” things like employers paying for Individual plans, (and using multiple accounts to get more usage) and aren’t yet enforcing the rules… but by the letter of the Anthropic terms, your employer should be paying Anthropic a whole lot more (and that’s one of the reasons why AI usage is going to get very very expensive as soon as the subsidies stop, you and a lot of other people are already paying a lot less than you should)
nl 16 hours ago [-]
AFAIK there is nothing in the ToS that forbids an employee paying for the 20x, $200/month plan. I could be wrong about this in which case it'd be useful to have a link the clause.
I think multiple plans are against the ToS, but I'm not doing that.
The Teams plans are more convenient for a number of reasons, but yes, they top out at the 6x plan, not the 20x plan.
I've re-read it and I'm pretty sure there is nothing that forbids a business paying for a 20x account. Notably they say this in the consumer ToS:
> If you use an email address owned by your employer or another organization, your Account may be linked to the organization's Anthropic enterprise account, and the organization’s administrator may be able to monitor and control the Account, including having access to Materials (defined below). We will provide notice to you before linking your Account to an organization's enterprise account.
which goes at least moderately close to indicating using it in a work environment is allowed.
andrekandre 16 hours ago [-]
i used to think the same as op but i think you are right. though they can change the eula at any time im sure....
skerit 11 hours ago [-]
> The individual plans are heavily subsidized
Not this again. There is not a single shred of evidence for this. In fact multiple times this year alone, people from Anthropic have said that inference and deployed models have positive margins. The big bucks are always being spent on training the next model.
nl 9 hours ago [-]
I agree the API prices are profitable (or at least break even) on an operational basis.
I don't think that means the subscriptions aren't subsidized though! I expect they are, at a carefully calibrated rate.
They want to maximize people trying them.
The real money is the enterprise plans which need a seat price plus API rates (unlike the Teams plans which are a pretty good deal).
queenkjuul 9 hours ago [-]
It costs $10B+ a quarter just to train?
It's a simple fact that paying enterprise token rates would cost many times what the individual plans cost. It may well be that the enterprise income outweighs the cheaper tokens on the individual plans for net profit, but using the individual accounts as subsidy is a common tech industry tactic and what you said doesn't prove they're not doing it, either.
It's a little hard to believe they wouldn't be trying to slow the burn rate for an IPO if the inference were truly so profitable. And i find anything they say a little hard to believe all the time anyway (though admittedly they're a lot more trustworthy than OAI)
jsisto 17 hours ago [-]
you sound like fun at parties
VladVladikoff 9 hours ago [-]
Anthropic killed it for me with all these silly security guardrails. As a programmer I use the models to check for vulnerabilities in my work. I also use the models for fun reverse engineering projects on the side. Neither of these activities are illegal. And I don’t appreciate being treated as a criminal. I asked Claude to do something the other day and it refused. I copied and pasted the exact same prompt into codex and it happily churned away at it for hours. What Anthropic has done to kill their own company is a bit sad.
browningstreet 3 hours ago [-]
They're expecting to IPO bigger than SpaceX later this year.
HN's take that they "kill their own company" seems wildly out of touch.
Truth is, Sam and Dario and Elon are all terrible and so are their orgs. Leaving any one of those companies for any of the others is wild.
semiquaver 24 hours ago [-]
My company still hasn’t been able to deploy wide access to Fable because it’s not available on a ZDR basis. This wasn’t mentioned in the article but I imagine this factor is not irrelevant.
rustcleaner 19 hours ago [-]
>ZDR
Zero Data Retention, for the uninitiated readers in this thread.
The no-ZDR is clearly to permit surveillance. I would be shocked if NSA wasn't all up in these SOTA model providers' systems.
Never subscribe!
NitpickLawyer 18 hours ago [-]
> The no-ZDR is clearly to permit surveillance.
It's to train better models. The three letter agencies don't need to spell it out in a ToS, they just access it if they want.
j_maffe 8 hours ago [-]
Not if there's a ZDR policy.
hirvi74 18 hours ago [-]
Ex-head of the NSA was on the board of directors for OpenAI for a while lol (and still might be there?).
Lucasoato 23 hours ago [-]
Same for me, they’re not offering it with the same data residency features of Opus, many enterprise companies can’t accept compromises.
kwisatzh 23 hours ago [-]
This is the biggest factor. FT (or the analysts it cites) fumbled the ball in this article.
janalsncm 22 hours ago [-]
> Data retention rules imposed by the Trump administration have also hampered Fable’s adoption, according to Ara Kharazian, chief economist at Ramp.
This is in the article.
semiquaver 21 hours ago [-]
It’s far from clear to me that this is directly connected to the no-ZDR requirement. I’m heavily involved in this stuff with my company and I’ve never heard that the lack of ZDR fable is a Trump admin thing.
topbanana 13 hours ago [-]
That's the reason at my shop. Real shame, as there's really no limit on token spend
jimmydoe 19 hours ago [-]
ZDR for fable is coming, very soon.
usaar333 22 hours ago [-]
It was in one paragraph mentioned, without much detail. But yes, I agree it is one of the largest barriers.
x3n0ph3n3 23 hours ago [-]
The same for my organization. It's explicitly forbidden in my organization because it's not available with ZDR.
Vespasian 17 hours ago [-]
I get where it's coming from but I think it is also very funny if someone actually believes their claims (except as a CYA strategy)
The AI companies have already shown how little they care about other peoples intellectual property and they won't care about their customers either if it stands between them and the promises they made to their investors.
semiquaver 10 hours ago [-]
This might be a reasonable thing for an individual consumer to believe about the value and enforceability of agreements with large companies, but given the context I think you’re off base.
ZDRs are backed by contracts negotiated against extremely capable and well-resourced counterparties. Being found to violate them would bankrupt even OpenAI and Anthropic many, many times over.
To the extent that you view these companies as your adversaries, it’s worth improving your mental model of their actual incentives and constraints.
Zigurd 23 hours ago [-]
For many coders including myself, LLM based coding agents work well enough to be useful, and in some cases worth paying for.
What I don't see is vast areas of industry finding $10s to $100s of billions of value in LLMs. There's no lint or compiler that can check for correctly constructed contracts. So LLMs, which should be useful to law firms, incur a lot more manual checking of their work than coding agents.
Less formal document production in other industries is likely to have less structure. That might not matter in some settings but I'm having trouble thinking of an example off the top of my head.
ray_kay777 21 hours ago [-]
While there's no linter for writing contracts, my experience (as a commercial lawyer) is that frontier LLMs are far better and error checking and far quicker at writing than the average senior lawyer. The main thing holding back further deployment (in my jurisdiction) are concerns around data residency, privilege and how fundamentally it will break an industry that is so heavily reliant on time based billing.
Using them for coding makes it easy to self check its work (assuming those pieces of work are "verifiable").
As proofreading will still need to happen, what do you think the appetite for lawyers is to do this kind of work? Do you think it will drive fees down significantly? Empower younger lawyers at firms who probably are the ones doing this checking for the partners? (Or will that just create a further divide).
I'm genuinely asking as I am not in law but all my family is and it's nice to see someone here that's thought about the impacts in that space.
ray_kay777 11 hours ago [-]
With access to good databases and tool use, hallucinations are basically a thing of the past (and very easy to verify). The hallucination issue typically comes from people using free ChatGPT or similar - tech usage and literacy is often poor and there's a lack of understanding of the risks of using something without proper legal database tools. Most lawyers would not have the faintest idea of what tool use in the context of an LLM even means. It's still seen as a magic box rather than a tool for real work.
Will it drive down fees significantly? Doubt it, there's not enough pressure on them, and the industry is resistant to change. Firms don't want to make a big deal about using it because clients will then ask why they're not getting a discount.
Will it empower younger lawyers? Not as much as I'd like. I'm very fortunate to have an employer that lets me use Claude Code for my work (to a limited extent). I think for 99.9% of lawyers it's not an option available to them, through a mix of concerns around AI usage and concerned IT departments. There could be great benefits, but it would rely on having to break out of traditional private practice which would make it difficult to get enough work.
I think in the near term LLMs will have much the same impact on the legal industry as it has on the software industry.
bostik 16 hours ago [-]
> concerns around data residency
Funny you should say that. I read somewhere (not on HN, but I think it was a post linked from here) that a number of law firms who deal with extremely sensitive documents have started buying amped-up Macbooks with 512GB of memory to be able to run local models.
These are businesses who literally - and for once this word fits - cannot afford to let some of those documents get anywhere outside their corporate walls.
senordevnyc 12 hours ago [-]
Macbook Pros max out at 128GB
bostik 6 hours ago [-]
Turns out the books do, studio setups can go up to 512GB.
There was a thread earlier this year about the option getting pulled from selection (https://news.ycombinator.com/item?id=47296302) but I'd guess you can still top the thing up yourself.
seer 17 hours ago [-]
But isn’t “writing a linter for contracts” the thing we need and probably would solve?
Coding agents are great because they have compilers, linters, test cases etc to ground themselves in.
With tools like OKF I’m sure most knowledge work would be distilled to its core data - it’s AST if you will, and then allow models to guard against hallucinations.
Checking if a case law exists is a tool call, you can demand provenance, it’s all _buildable_.
Hallucinations are “solvable” this way, so the rest is just time and adoption…
nprateem 19 hours ago [-]
Trouble is they're unreliable.
I compared insurance quotes last week. Needed cover for 2 brands, 1 company. Opus jumps up and down saying both brands need listing on the policy schedule. Human broker said not.
I told opus and it's the usual "thanks you're right" bollocks because it bothered to read in more detail and found that all business activities are covered.
algo_trader 15 hours ago [-]
>my experience (as a commercial lawyer)
How do you think the industry will respond to contract clause slop ??
I can foresee each side inserting 100s of innocent clauses with minute dependencies that provide hidden advanatges
ray_kay777 12 hours ago [-]
Good question - but the state of commercial contracts (pre-AI slop) is horrifically dysfunctional, inefficient, poorly thought through and inaccessible to the average person, so the bar is quite low. I think there will be some new malicious behaviours like that to keep an eye out for but I haven't seen them in practice (yet).
_joel 23 hours ago [-]
> There's no lint or compiler that can check for correctly constructed contracts.
I have a family member that is an attorney in housing law. She claimed that LLMs are not particularly useful for her work. If she asked a simple question like, "Find all the <insert specific housing laws> for all 50 states," then she still has to go and check every single one of the laws. Since the legislature is modified so often, she cannot look at, say, Maryland's law and know if it the LLM output was the 1990, 2014, 2018, or 2026 version of the codified law. In order to fact check the law, she has to look it up, and by that point in time, she has the answer she did the work of the LLM.
nrmitchi 16 hours ago [-]
This doesn’t sound as much that “LLMs can’t help me” as much as “trying to one-shot a complicated answer in a chatbot can’t help me”. Output straight from a generative model definitely should not be trusted.
But that doesn’t mean that an appropriate harness (even Claude code and a dynamic workflow) can’t complete, and validate, an answer.
suddenlybananas 12 hours ago [-]
Seems like it would be a massive market to make such a harness, and yet you're making it sound very easy to do. It would be surprising no one has done it if it were.
nrmitchi 7 hours ago [-]
Harvey and Clio both make entire products around this. They have raised billions of dollars.
That's not really the point though; it doesn't have to be "legal specific".
The key is in understanding that "Basic chatbots (claude, chatgpt, etc) hallucinate. Independently verify the output (even if it's another model doing the verification) before assuming it's correct".
nottorp 16 hours ago [-]
You can reproduce these results even if you're not a lawyer, if you've dabbled in any MMO that has been around for a few years and had multiple patches.
Ask any LLM about a random game mechanic and then be prepared to verify if the answer is for the patch in 2025, 2023 or 2020...
ipaddr 18 hours ago [-]
This is my experience if you need the data to be accurate often it isn't and you only know that because you had to check because it was important. Things not important you never check.
mtrovo 12 hours ago [-]
LLMs are not magic boxes, they hallucinate and deviate quite a lot when asked hard to verify questions.
The only way we got a head start of using it for coding and maths was to have some formal method of validating the output as part of the training and inference time.
carlosjobim 22 hours ago [-]
> What I don't see is vast areas of industry finding $10s to $100s of billions of value in LLMs.
Translation between languages.
That value dwarfs all programming value that can be had. Economically, culturally, scientifically, spiritually.
nisiddharth 18 hours ago [-]
A model bundled in your phone can do the translation, you're thinking too much of it.
carlosjobim 12 hours ago [-]
A model bundled in your phone cannot do commercial grade quality of translations.
nitwit005 21 hours ago [-]
Between human languages? The problem will be Google is giving that away for free.
theshrike79 15 hours ago [-]
You can do that with single words or phrases.
Now do a full movie's subtitles with google translate, in an automated fashion. So that it understands the context and still translates it correctly.
carlosjobim 21 hours ago [-]
The quality of Google Translate is so low that it isn't even in the picture.
bluegatty 20 hours ago [-]
Social issues aside, that's a tiny slice of the economics of AI.
Just by nature of how often it's used etc.
You need workloads for AI be cost effective: software, automation etc.
carlosjobim 20 hours ago [-]
International trade is a giant sector already, and potentially much bigger than that, now that LLM assisted translation makes it easy to offer your products and services in any language, or make complicated and sensitive deals without a common language.
Hackers always rage and down vote every time I mention this, because they are unable to see beyond their small world. Why didn't they learn that their part of the internet is 0,000000001% of what the world uses the internet for today. It's going to be the same with LLMs. Programming and hacker stuff is going to be 0,0000000000000000000000000000001% of what the world uses AI for. But translation is going to be in the top 5 of use cases.
bluegatty 18 hours ago [-]
Product marketing and contracts will be a 'common use case' we think of, but it will be 0.000001% of tokens consumed.
A developer using sub-agents will consume more tokens 1 Day than a marketing manager will consume 1 Month, easily.
Unless there is something inherently automated about the nature of the AI, it will be a tiny % use case.
Even a lawyer, using AI daily for contracts - that will be relatively light use. They'll make more use doing legal research etc.
Developers and Automation are the 'primary' uses cases for AI, and in the future, we'll start to see AI integrated into Apps - that will be 85% of tokens consumed.
Yes - once translation becomes realtime, and we have our Star Trek Universal Translators, then translation will become more visible, but even by then, a relatively small part of overall consumption, even if it's more highly visible.
carlosjobim 11 hours ago [-]
If a salesperson uses 100 000 tokes for translation, which gives him $10 000 in sales.
And a programmer uses one million tokes, which gives him $100 in sales (or value).
What is then the value of a token?
The comment I answered asked where there is an industry finding 10s or 100s of billion of dollars in value from LLMs. The answer is translation. It's the value they as customers get out of the LLMs, not what cost they are paying for the LLMs.
Value has to be counted in production, not consumption.
> Developers and Automation are the 'primary' uses cases for AI
Just like programmers and scientists were the primary users of the Internet when it began. But things change rapidly.
thinking_cactus 19 hours ago [-]
I think quite cheap models will probably work with translation (I think for many use cases Google's now fairly old tech works well enough and is free?).
The main problem is I think you're assuming because the current translation market is large (I'm just going to assume it's ~100B in size just from a cursory search), then it will remain large with LLMs. If LLMs are much cheaper than humans, even with a lot of growth in translation volume the total spend may not compensate for it (again most of the volume will probably be using almost free models?). Another is assuming that because something is valuable you can charge a lot for it. Like, oxygen from air is extremely valuable to us. If oxygen somehow depleted we would die almost instantly. It does not mean everyone goes around purchasing oxygen or even less that you can charge absurd amounts for it.
carlosjobim 11 hours ago [-]
LLMs are creating a new translation market.
Before AI translation became available, I would hire professional translators. Their rate was about $50-100 for a detailed product page into one language, and it would take them a few days to deliver.
Now with AI translation, I pay about $120 per year for unlimited translation. Meaning dozens of product pages into 5 languages, plus e-mail back and forth with hundreds of customers. All instantly at my convenience.
What this means is that a whole lot of people, sectors and businesses who would never hire professional translator can now have high quality translation at their disposal for a cheap price.
> Another is assuming that because something is valuable you can charge a lot for it.
Even if LLM translation won't deliver trillions of dollars in income to the AI companies, it will without a doubt deliver trillions of dollars in value to customers and users within the coming few years.
Good human translators are still higher in quality than any AI and will always be. But AI translation is currently far beyond good enough for all use cases, except fine literature and maybe complicated juridical stuff. But I'm not familiar with those sectors.
fullshark 19 hours ago [-]
I think OP's comment is about volume/scale more so than practical application.
eikenberry 21 hours ago [-]
But translations don't require anything close to SOTA level models. Translations will be high volume, low margin transactions. That will not save Anthropic.
carlosjobim 21 hours ago [-]
Maybe not Anthropic, but LLM translations has a value counted in the trillions of dollars easily.
tock 19 hours ago [-]
I see it's important but trillions? How did you quantify that number?
carlosjobim 12 hours ago [-]
Tourism and travel is 10% of global GDP, or about 10 trillion dollars. And it's not the only sector where translation between languages is a benefit. However, I am not sure that the trillions of dollars of value from LLM translation in the coming few years will translate into trillions of dollars of income for the AI companies.
tock 9 hours ago [-]
Yep translation is going to be a free feature on all devices using on-device LLMs. I see no revenue flowing to AI companies.
carlosjobim 9 hours ago [-]
Commercial grade translation between dozens of languages would take enormous amounts of disk space.
But we'll see. Text editing tools are included on all digital devices, yet companies pay for commercial solutions like MS Office. Cameras and basic image editing tools are included on all smart phones, yet there are millions of advertising agencies around the world. And so on.
tock 8 hours ago [-]
LLM frontier companies are more akin to infrastructure like chips and airplanes and the internet than applications built on top of it. There will be products built on top of them that generate revenue.
eg. processors are needed for everything on the planet. No Intel isn't going to generate annual revenue in the trillions because of that. In the case of frontier companies the moat is even smaller with dozens of players competing. LLMs will become a low-cost commodity. You meanwhile can make billions building applications using them though.
dlcarrier 1 days ago [-]
Transformer models are quickly becoming a commodity, and I suspect in time we'll all be running them locally. Even now, you can run something pretty useful on a 16 GB graphics card, and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
rustcleaner 19 hours ago [-]
>and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
I doubt it. The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription. Even if my #NeverSubscribe movement succeeds, all that misdirected hardware [into datacenters] won't likely be practical for home use.
dlcarrier 16 hours ago [-]
I had an internship at HP in the early 2000's, and my computer was a terminal that connected to an X server in a data center. Doing everything as a service is not a new idea, and suppliers have been trying to push it for decades.
The only time it really succeeds is when regulations force it. I can store files on my own computer or a home NAS just fine, but if I start a medical practice, I have pretty much no hope of being HIPAA compliant without signing up for a data hosting service. The same goes for tax preparation, banking, and several other fields.
It may very well be the case in China, Europe, or Australia that regulatory restrictions force users into models-as-a-service, but in the US, regulations censoring models, even if they frame that censorship as a safety measure, won't pass constitutional muster.
curious_cat_163 19 hours ago [-]
> The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription.
I had not considered that as a possibility. It is somewhat dark and unlikely, imho, but a real possibility.
I am more inclined to believe that the fab capacity will grow over time and "commodity" compute will be available to us all again.
However, I can also imagine going back to the 60s era "hyper verticalized" mainframes, in which case, the frontier labs might not just be producing models, but chips and an ecosystem around themselves and their suppliers/customers.
queenkjuul 9 hours ago [-]
But the Chinese models exist, are competitive, and have locally useful quantized versions.
If you're actually willing to pick up old DC inference hardware on the cheap you can already do a lot at home. The new hardware not so much admittedly but if the Chinese models keep getting better and running on less hardware... I don't see a reason to be quite so pessimistic.
chr15m 23 hours ago [-]
Yep local models will be good enough for most things you need to do, in the same way as most people need a laptop not a supercomputer.
DiscourseFan 21 hours ago [-]
Well nobody thought you could write software that helps tightly coordinate processes happening simultaneously from millions upon millions of nodes on nearly every corner of the world, but here we are. When industry expands further off earth, we will need more complex and intricate software to coordinate its movements, why wouldn’t our systems become more powerful. If you are hopeful for humanity than you must expect the scale of industrial necessity to only ever increase alongside the imagination and capacity of its people’s.
chr15m 14 hours ago [-]
I think the opposite can be true. Billions of humans could be happy on the Earth we have now for a billion years.
gritzko 16 hours ago [-]
Also, pairing a wireless mouse to a laptop is not getting easier as years pass.
FeepingCreature 13 hours ago [-]
This won't happen so long as the bandwidth bottleneck is a thing. Local models are just several orders less efficient at scale, and this limitation is inherent to the architecture and won't go away unless both hardware and model structure change enormously. There are really only two usecases I can see for local models going forward:
- 10-1000 person groups such as corporations where you can amortize serious hardware with parallel use
- Porn.
nate 22 hours ago [-]
Of course I echo the universal: Opus 5 sucks. But what also sucks is still using 4.8. Are you all seeing this? It's like the older models got dumber just before Fable and Opus 5 were coming out? I've heard the theory that it's because 4.8 is getting put on older hardware? And I imagine that just ratchets down reasoning time then possibly?
Sometimes I'm just trying to sus out if I'm truly seeing things these days or going a little nuts :)
purpleidea 18 hours ago [-]
I have noticed very serious degradation in performance from Anthropic. I've switched away from Opus 5, but 4.8 is still much worse than how it was before Fable came out. It's constantly making what seems like obvious mistakes... I point them out, it's constantly apologizing.
I don't know why...
Is it A/B testing?
Is it load shedding?
Is it because I'm in Canada?
Is it because I'm not on the Claude Max plan?
Is it because I'm not paying via API?
Is it because I'm not paying via Bedrock?
Is it because the U.S. is worried people are distilling?
Is it because the U.S. wants to keep the top capability to themselves?
I think open models are the future. Anthropic is killing their reputation so fast. If they don't come clean I think they're cooked.
somenameforme 21 hours ago [-]
This is yet another reason why I think local models will win in the future. They're almost certainly A/B testing all sorts of opaque stuff that people have no clue about, hence the various 'How's Claude doing this session?' popups.
So what you are paying for may vary on a day by day basis, which is quite undesirable, even if their main goal is simply to make a better model. When it comes to a tool, I'd rather have consistent mediocrity than instability.
zsoltkacsandi 18 hours ago [-]
Yes. OpenAI is doing this as well, but they are much more transparent about it.
samrj12 21 hours ago [-]
I definitely agree. I honestly cannot do tasks that require even minor complexity. Opus 5 keeps forgetting things in context as well and coding conventions. Really cannot build with CC without Fable.
zsoltkacsandi 18 hours ago [-]
Favourite conversation with Opus 5:
Me: why did you add HTMX?
Opus 5: you asked it twice.
Me: quote the exact sentence(s) where I asked it.
Opus 5: I can’t because you didn’t.
Gareth321 14 hours ago [-]
RIGHT!? The data suggests this isn't true but every fibre of my being is convinced that Opus 5 High was EXCELLENT at launch and has been lobotomised since then. My benchmark is Sol High. I've been using both consistently and either Sol High suddenly became MUCH more capable - and the data does not support that - or Opus 5 became much dumber. It's so bad I can't even use it anymore.
drbscl 14 hours ago [-]
IMO Opus 5 wasn't that great at launch, but 4.8 was definitely downgraded just prior.
Agree that Sol is a great model. I find that it's improved a bit but I mostly attribute this to its eagerness to use the harness' memory features. (I'm using it in Hermes Agent, FWIW)
gandreani 21 hours ago [-]
I'm not an expert by any means whatsoever but deploying models is not a straightforward task. There's a lot of levers to pull and I bet when models get "downgraded" to older hardware they do so WITHOUT the same stringent quality control of the output as they do when they release it.
I don't think it's something deliberately malicious like planned obsolescence but it's more like startup culture of "just make it fit in this sprint".
Gareth321 14 hours ago [-]
I attribute intent. These are for-profit companies with more leveraged capex than any other industry in history. They have enormous pressure to optimise their limited compute. Especially Anthropic, which is far more limited than OpenAI. Of course they're pulling those levers in the background to achieve "good enough." If they can reduce memory consumption by 20% and their metrics show a 4% loss of intelligence, they might very well pull that lever. These decisions compound.
Your explanation is probably the most likely and largest contributor. Anthropic states that Opsu 5 is a "pinned snapshot". They claim the weights and model configuration are not silently updated, BUT the surrounding serving infrastructure can change, including the request router, safety classifiers, and sampling logic. Anthropic has stated that if behaviour unexpectedly changes on a stable model ID, an infrastructure update is the most likely cause.
Further, "High" isn't a fixed amount of compute. Anthropic describes effort as a "behavioural signal" and not a token budget, with the model deciding how much thinking to do. So their "High" might be "Low" now, and we would never know.
Finally, I strongly suspect some quantisation or KV-cache compression is happening. Anthropic doesn't clearly delineate whether this would fall under the pinned weights and configuration, or the infrastructure, which almost certainly guarantees it's the latter. Forgetting earlier information, poor retrieval of details, contradicting previous conclusions, hallucination, degraded instruction-following, and losing the thread during complicated tasks are all symptoms of quantisation and compression.
make3 21 hours ago [-]
For me, Opus 5 mostly sucked because of its incomprehensible writing style. Having the system prompt focus on writing in terms that are easier to understand helped a bit
madrox 17 hours ago [-]
Anthropic lost all good will with me. Everything from their policies to their rhetoric has an air of "we don't trust you." The way they treated users who wanted to use them for OpenClaw didn't sit well with me, and then the Fable nonsense was the last straw.
And they're somehow shocked users aren't loyal to that. It isn't about cost.
theshrike79 15 hours ago [-]
Which AI companies still have your good will?
madrox 3 hours ago [-]
In addition to what another reply has said, I'll just say any company that does not make me feel like they'll ban me if I hold their model wrong.
The bar is low.
nullbio 12 hours ago [-]
OpenAI is actually great from a users perspective. Responsive on Github issues, engage with the community, listen to feedback, constantly give subscription resets, fair subscription rates, speak out against fearmongering rhetoric, not cutting off your workflow mid task if your subscription runs out, and the list goes on.
kiratp 23 hours ago [-]
The actual issue, is suspect, is that Anthropic won’t provide ZDR for Fable. Makes it a non started for a large percentage of businesses.
nullbio 12 hours ago [-]
Yeah, and it's no secret why they won't. Your data is the only reason they even still provide subscriptions.
ericol 1 days ago [-]
The problem is that they are not solving the problem they ought to be solving.
I don't want Shakespeare, I want Bob the builder.
Half of my work is telling claude how to behave. I'm pretty certain they have enough _data_ to realize people do the same thing time and time again.
They already do, extensively. Its not Shakespeare, in fact, it sucks at prose and creativity, like most newer llms.
But people are not as alike as you think. I doubt I share your unique preferences.
That said, I don't spend much time telling it how to behave. Are you sure you're not fighting the default system prompt?
ericol 21 hours ago [-]
[dead]
rowanG077 24 hours ago [-]
This is really it for me as well. At the heights of complexity AI can do magical things. But really a lot of the time I just want it to do mundane things right. And currently it just cannot. It writes garbage text, consistently ignores something you have told it, makes mistakes a human makes once but the AI remains uncorrectable.
Barrin92 22 hours ago [-]
>Half of my work is telling claude how to behave
Dijkstra in the Foolishness of Natural Language Programming
[...] the "naturalness" with which we use our native tongues boils down to the ease with which we can use them for making statements the nonsense of which is not obvious. It may be illuminating to try to imagine what would have happened if, right from the start our native tongue would have been the only vehicle for the input into and the output from our information processing equipment. My considered guess is that history would, in a sense, have repeated itself, and that computer science would consist mainly of the indeed black art how to bootstrap from there to a sufficiently well-defined formal system.
Imagine if only we had languages at our fingertips whose explicit purpose was to precisely and unambiguously tell a machine what to do!
I've spent many years doing work in compliance for airlines. Hundreds of pages of documentation to describe the different rules and regulations which must be followed in specific scenarios. We'd convert those documents into a programmatic definition (rule engine) to alert when rules are at risk of not being followed.
This work was fraught with bugs, a large portion of which came down to disagreement of what was coded v/s what was written. Even if you had airlines sign off sentence by sentence exactly what you wrote down in English, that's too open to interpretation.
People don't appreciate how the same sentence can be read five different ways by different people (or the same person on different days). We had to structure our documents to be closer to pseudo-code than to English to get any meaningful consensus on the definition.
nottorp 16 hours ago [-]
> Imagine if only we had languages at our fingertips whose explicit purpose was to precisely and unambiguously tell a machine what to do!
But that would require people to think! And the marketing is they don't need to do that...
Last consumer (ish!) product that required people to learn something to use it was PalmOS with Grafitti, wasn't it?
YuechenLi 1 days ago [-]
Fable is just way too expensive and limited compared to GPT 5.6 Sol, and the only task that requires that level of intelligence is frontier scientific research. I use GPT/Codex primarily for coding and usually keep Claude on Sonnet 5 most of the time as I use Claude primarily to debug/brainstorm/make frontends as a supplement to GPT.
jml78 1 days ago [-]
I do not agree. Fable is the only model I can leave unattended and give me results for part of my work which is just devops related tasking.
I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.
I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic
gizmodo59 24 hours ago [-]
I second the parent comment. 5.6 sol xhigh is not only better than fable I can also run it forever without worrying about limits. The frontend has gotten much better too with the plugins that come with codex.
mapontosevenths 23 hours ago [-]
Agreed. I have the highest individual plan for both. I run out of Fable credits midweek, while I usually have some credit with Chatgpt left despite it having to carry Fables load for the second half of the week.
Also, Sol doesn't refuse constantly and speaks like an engineer rather than a deranged academic.
Opus 5 is legitimately terrible and can't or won't follow instructions. It is of negative utility and does more harm than good to my codebase.
bluegatty 20 hours ago [-]
OP5 will do long running work, you just have to make it write a plan.
And - trick - give it a little cli so it can run Codex if you have them both.
Let it do codex to do the bulk of the work, get a OP5 sub-agent to audit the work of the codex worker.
Just let Op5 manage and have 'specific oversight.
You can run for 2 days on 1 context window in the manager, the advantage is that it will stick to a broad plan.
mapontosevenths 7 hours ago [-]
It's kind of funny you describe it that way. I'm literally doing the exact opposite, which is why I have both.
I have Sol xhigh drive Claude via tmux and I get amazing results until I run out of Fable. Then Opus comes in and starts acting like some sort of autistic academic with OCD.
It tries as hard as Fable, but isn't smart enough to do it well. It starts designing ever more elaborate tests, frameworks, and procedures while making up rules for itself and piling them on top of each other until nothing gets done. It's the ultimate bureaucrat.
Worse - More than once it's spent days in a loop because it invented constraints for itself that it couldn't satisfy then lied and told Sol that the user imposed those limitations. I don't know if it's actively avoiding real work, or just isn't capable enough to work the guardrails that were obviously forced into it.
gizmodo59 22 hours ago [-]
I’m referring to Fable vs 5.6 Sol. Opus 5 being bad is universal at this point.
hodgehog11 23 hours ago [-]
In a frontier scientific research environment, funding is often limited, so personal subscriptions are more common. Fable can hit a 5-hour usage limit on the Max subscription tier before it finishes a single complex math prompt. Most of the time you get one or two if you reduce context and question complexity. Most of my colleagues are cancelling their subscriptions because of that, and just using Sol instead, which is virtually unlimited on the Pro tier.
YuechenLi 22 hours ago [-]
Hmm. Would you mind sharing an example of a complex math prompt that you would use? Because I found that even Sonnet can solve fairly complex math problems fairly easily if you give it the right tools, so I'd like to give it a shot myself if you'd like.
hodgehog11 20 hours ago [-]
[dead]
redox99 19 hours ago [-]
> the only task that requires that level of intelligence is frontier scientific research
Even fable is way below human level at many tasks. If fable were as good as you claim, all computer jobs other than "frontier scientific research" would have been replaced by now.
hellisothers 1 days ago [-]
I thought Sol was on par with Opus, so comparing it to Fable is apples and (very expensive) oranges?
rybosworld 1 days ago [-]
I've used all three extensively.
Most of the benchmarks have exceeded their usefulness. Opus 5 beats fable 5 on many of them. Anyone who has used both models will notice immediately that this doesn't translate to the real world. Opus 5 is nothing short of a regression from Opus 4.8. Fable is genuinely a great model so long as you don't trigger a guard rail and it downgrades.
Sol in my experience isn't significantly different than fable ignoring that Sol burns usage 10x faster but the end result is hard to differentiate.
GLM 5.3 is a hair behind these two.
An anecdote but not an original one from the people I talk to.
rafaelmn 1 days ago [-]
Sol routinely catches stuff in review that fable misses for me. It's impossible to compare them meaningfully because it's a complete dice roll - but in practice using both in my projects I can get work done with both, and Opus 5 is far more tedious.
But Fable security false positives and pricing just make it not worth compared to Sol IMO.
YuechenLi 1 days ago [-]
Yeah, it's pretty much apples to oranges, and I don't consider GPT and Claude to be interchangeable at all. From my anecdotal experience, GPTs generally codes more creatively and verbosely but Claudes tend to code more carefully and precisely, so the result is that GPTs generally finds more creative solutions to problems but also writes buggier code, which is why I converged on the setup of GPT/Codex for implementation and Claude for debugging, which feels more like a force multiplier than using each model individually.
markbao 1 days ago [-]
I’m not sure if the underlying data is counting subscription use for Fable, which is where a lot of people are using it because token pricing is very expensive. I wouldn’t be surprised if this was counting enterprise token usage only. As rich as enterprise customers are, they’re not exactly willing to double the cost of SWE salaries on tokens.
Either way, a model used to solve the top 10% of problems that people use AI to solve for, being used 10% of the time … seems like it’s in a decent place.
I find Fable indispensable, and measurably better than alternative models, for complex feature development in an existing codebase. It’s the closest thing I’ve seen to nearly one-shotting features. Even still, I only use it for the hardest features and Opus 5 does a good enough job on the rest.
wj 23 hours ago [-]
Very much agree on Fable. Over the past month or so it has shown to be the only Anthropic model that can understand a largish dbt model codebase. Opus 5 gets almost everything wrong. (High reasoning on both)
dmix 23 hours ago [-]
More importantly other way cheaper models can do what Opus 5 does. So you can pay for Claude to use Fable 5 exclusively for harder stuff and planning, then get the same value you'd otherwise get from switching back to Opus by using other cheap LLMs for day-to-day coding tasks.
brianwawok 23 hours ago [-]
Wouldn’t you need to bump opus reasoning a few levels to be apples to apples?
anon7000 24 hours ago [-]
Yeah as soon as CFOs realized AI was racing to become one of the most expensive line items along with salaries and AWS bills, they started cracking down on the most expensive ones.
Roark66 9 hours ago [-]
You know what. As a big Anthropic fan that pays €120 a month for Claude Max x5 and maybe €30 a month for API access I got extremely annoyed at the extremely variable quality of service I'm getting.
It is not only that every single turn with opus on high reasoning (high is the middle setting) takes at least 5-7min. It is barely usable interactively. Instead of a chat it feels like you're sending emails to it. Tasks that used to take an hour when it "reasoned" for 45s before it started doing anything now take almost entire day.
At least until few weeks ago it was horribly, mind bogglingly slow (a little better during US nighttime), but the quality was still good. I could not do things interactively, but providing prompts were fine it built stuff fine.
This is no longer the case. It makes stupid errors all the time. So you cannot leave it to complete some work, for example write infrastructure migration scripts a night before then you simply run the scripts and perform the migration during the day. Nope, every single script has stupid issues requiring use of the model to fix them. As they are written in it's own "spaghetti code" fixing them by hand is not an option.
It is clear to me they are doing some shenanigans behind the scenes to try and optimise their compute use. Either they quantized these models dynamically or do other things that affect quality.
In top of that they now do this stupid fingerprinting.
Anyone who knows how output vectors are turned into tokens knows it will eat up a lot of compute or destroy quality.
hbarka 1 days ago [-]
Anthropic’s issue is churn because of the peak verbosity vomit coming out of Opus 5/Mythos/Fable. What the hell did they train it on. The sane model is still Opus 4.6.
- Opus 4.6 was the last Opus generation that got a lot of use by Anthropic's own employees
- After that they primarily used Mythos internally
- 4.7, 4.8 and 5 were RLAIFd by Mythos "teachers"
- Hence why 4.6 is the last Opus gen who doesn't report back like a robot wanting to cover every potential hole another AI system would've spotted and criticized
- Hence why coding style in Opus 5 also gets criticized, not only behavior in CC
bitexploder 23 hours ago [-]
Commented elsewhere, I still use Opus 4.6 because it is the only model that feels decent to interact with. 4.8 is decent and some times smarter but you can see it trending towards Opus 5 levels of nonsense. I use Opus 5 when I don't need to interact. Fable or Opus 4.6 are the only Anthropic models I like interacting with ATM.
nl 18 hours ago [-]
I've retweeted and posted this here before too. I think there is probably some truth to it.
Worth noting that Fable (ie, Mythos) is actually nice to interact with.
ericol 1 days ago [-]
There's actually a tool called vomit [1] of all names to fix exactly what what is being discussed here.
Perhaps there is a sophisticated subtle poisoning attack that makes models behave like that?
rglover 1 days ago [-]
Karma backed over their dogma.
I switched from Opus 4.7/4.8 to test Kimi K3 a few weeks back and the test hasn't finished; it's my daily driver now.
Given their general behavior and preference toward social engineering to scare the shit out of normal people...this seems fitting.
maverick98 8 hours ago [-]
Personally I've been turned off by Anthropic's stance in many topics, regarding replacing software engineers and other professions, shady pricing, being on a high horse etc. In my work I use their models when I absolutely have to, otherwise I use cheap models to do my work.
gwbas1c 3 hours ago [-]
I just don't like Anthropic's 5.0 models. When I tried Fable and Opus 5.0, they were super-slow and gave me sub-par results. I went back to Opus 4.8, and then switched to GPT 5.6 Luna. It's cheaper and better than Fable.
This is just a case of competition working: Fable (and Opus 5.0) just aren't as good as Anthropic believes, and there's no switching cost, so everyone's moving away.
runako 22 hours ago [-]
The US government basically told them they can't sell "Fable" and so they aren't. That's probably 90% of the story.
Once the government stepped in, their 5th generation was effectively killed. They either had to lock it up (Mythos), neuter it (Fable), or leave people with the perception they were overpaying for a weaker model (Opus 5).
Hopefully Anthropic has learned to anticipate this risk and has a plan for rollout of their next model that plans for capricious ad-hoc regulation.
rustcleaner 19 hours ago [-]
>They either had to lock it up (Mythos), neuter it (Fable), or leave people with the perception they were overpaying for a weaker model (Opus 5).
Option C: they could "get hacked and the models exfiltrated" plausibly. Then the horses will be out and there would be no more point for the USGov's barn door.
somenameforme 21 hours ago [-]
What did they think was going to happen when they were actively doomsday fearmongering around the release of their own model? I assume somehow they thought this would lead to a moat with them being left safely inside the castle, but this response, and Reagan's 9 Words, were always infinitely more likely. It was a demonstration of a child-like level understanding of how regulatory capture works.
runako 20 hours ago [-]
The PR from the US frontier labs has been poor to terrible from the beginning.
Among other things, "we have stolen the collected works of your culture, now help us grow so that you can all lose your livelihoods and become our serfs" has to be one of the absolute worst marketing approaches in history.
rustcleaner 19 hours ago [-]
>so that you can all lose your livelihoods and become our serfs
Crimemaxxing about to go off the charts!
ComplexSystems 21 hours ago [-]
The government seems like it's in a rough spot. If they let Mythos out, they seem worried people could use it to mass-hack the internet. China seems to not care so much about this and they're right behind. I don't really know what the answer is.
dannyw 20 hours ago [-]
I wonder how much of it is fears over "mass-hacking the internet", and how much of it is fears over the model discovering various NSA/CIA/etc "tailored access operations", and other deliberate side-channels / vulnerabilities / etc?
(Or perhaps vulns that NSA/etc discovered and has been keeping it private; as they're known to do).
I feel there's a lot that's unaccounted for, and the whole "AWS team reports a 'jailbreak' that is just 'review this codebase'" story doesn't add up.
I wonder if there were some parallel construction going on, and if at the same time, the NSA started losing the exploits they had because it was getting patched.
Revanche1367 19 hours ago [-]
I think the latter is likely their primary concern. Certain models making hackers more effective doesn’t change that black hat hacking will still be illegal and that’s always been the prime deterrent against capable hackers.
While there is some fallout likely with much more effective hacking being easily accessible, I’m sure govt analysts (unless they were let go) have their own prediction models telling them it’s inevitable that this technology eventually makes it to everyone they don’t want having it, what with China seemingly releasing every progress they make openly. Which makes me think that they’re preparing for that inevitability by hardening the govt systems currently in place and/or by burying the secrets they want to keep hidden deeper underground.
Given the history of the US and this particular administration, I feel burying things deeper is a greater priority.
rustcleaner 19 hours ago [-]
One thing is for sure, the USA certainly aren't the good guys anymore (if they ever really were).
-t. American
alightsoul 20 hours ago [-]
China uses the strategy of letting dangerous technologies loose which causes disruption in the short term but makes people do the right thing like secure their software. The US by comparison gives me the impression of wanting to leave the internet vulnerable by not making Mythos public so that only the US government can use Mythos to gain access to whatever system they like, which is the same thing the pegasus software does
runako 20 hours ago [-]
Whether or not the government chooses to regulate the space, capricious after-the-fact regulation is the worst of all possible worlds. The ~equivalent models from OpenAI did not get the same treatment (favoritism?), and it's not clear the government has produced even rough guidelines about how to be compliant going forward.
Model prep costs far too much money to operate under this kind of regulatory regime.
0xbadcafebee 16 hours ago [-]
The government isn't in a rough spot, they're just idiots. Chinese models already match Mythos in red-teaming. Anyone can use Chinese models any time they want. By holding people back from Mythos they're pushing people to Chinese models. And it worked: Chinese models now make up 70% of OpenRouter tokens (complete inversion from a year ago). Even if Chinese models weren't that great, anyone can fine-tune an open weight model specifically for red-teaming and it'll outperform any other model. So this was always going to happen.
A government staffed by logical, sane people would have figured out how not to encourage this, like mobilizing the IT and security industries to secure their products faster, while also cracking down on the "our AI is the most dangerous tool in the world" rhetoric bandied about by the frontier companies. But we don't have that, we have a government of Loony Tunes characters. When you put extremists in office you get extremist behavior.
nl 19 hours ago [-]
> The US government basically told them they can't sell "Fable" and so they aren't.... They either had to lock it up (Mythos), neuter it (Fable),
But this isn't what happened? It was only a couple of months ago! How can we be getting this completely backwards already!
Fable was aready what they were selling to the general public, and that is what the US government stopped them selling (not Mythos!)
Anthropic tuned the already existing classifier to handle the use case the government highlighted and then they went back to selling it.
They actually loosened the restrictions on ML-programming using it too.
I use Fable as much as I can and I've never had a cyber refusal for it.
runako 18 hours ago [-]
You're agreeing with me.
Mythos was locked down to a few customers.
Fable went wide.
USG told them to stop selling Fable.
They tweaked Fable to appease the USG, but it clearly refuses/downshifts more requests than e.g. Opus 5. (I have run into Fable refusals even doing really normal stuff like fixing bugs reported by static security scanners like Brakeman. Not downshifts, outright refusals.)
So the net result is they put "Fable" back on the market, but it was perceived to be a worse overall experience, at much higher cost. (Remember all this happened before most users had really even had a chance to put Fable through its paces.)
That's where we are today: Mythos is locked, Fable is neutered. They probably can't fix this until they are ready to release something that's not called Fable, or that they can say has some key architectural differences to Fable. The sword of capricious regulation is always going to be hanging over Fable.
nl 17 hours ago [-]
> The US government basically told them they can't sell "Fable" and so they aren't.
> You're agreeing with me.
No I'm not. They are selling it.
I used Fable very extensively before the ban (100% usage in every window available) and continue to use it now.
I haven't noticed any change in the refusals before vs after.
> it was perceived to be a worse overall experience,
Sure, people will believe whatever they want to believe. Doesn't make it right though.
Zigurd 10 hours ago [-]
People promoting investment in AI are fooling you with bad TAM estimates. For example, if we valued every spreadsheet created using the same metrics of pre-automation paper spreadsheets, spreadsheets would be a $100 trillion business.
Computing technologies are relentlessly deflationary. If the value of their TAM wasn't a fraction of a manual process they replace, they wouldn't have a productivity advantage. And the amount of TAM per unit often declines over the life of that product category.
I would be unsurprised to find investment in data centers to be 10X what was really needed. The same goes for where the LLM S curve starts to flatten.
383848484848 10 hours ago [-]
in other news water is wet
this is about selling credit to clueless boomers
syntaxing 24 hours ago [-]
GLM series has made it very practical to self host. If the new update for Deepseek flash holds up, I think it would be silly for some companies to not self host.
polski-g 23 hours ago [-]
We're spending 225k a year on tokens. No reason not to buy the hardware necessary to run DS4 at this point.
nicce 15 hours ago [-]
Are you saying that you already spend so much money for tokens that too late to buy the hardware? Sounds like a trap...
syntaxing 22 hours ago [-]
Agreed, and you can write cool infra agents to do stuff for you that runs during off hours like nightly tests and triage.
fxtentacle 1 days ago [-]
In my opinion, the big issue with Fable is that Claude Code cannot use it properly. I know, that sounds weird, but I've had Fable run down the wrong lane (and never stop) or give up and claim that something was impossible so many times (until I pointed at a GitHub repo that solves the "impossible" issue).
But a while ago, I had access to a Fable harness that just never gives up. And that verifies itself. It burned $100 in API tokens in 15 minutes ... but it succeeded for all the prompts where Fable + Claude had failed.
And I believe that's a real issue for Anthropic. Fable+Claude is not too expensive thanks to the subscription, but Claude severely nerfs Fable. To save money, I guess. Fable API + Custom Harness is a different class, it's so much better. But API tokens are so expensive, you're cheaper off hiring a freelancer.
ipnon 24 hours ago [-]
It kind of has the personality of those students that get so stuck on one promising idea they lose sight of the problem.
ieie3366 1 days ago [-]
Fable is not a tool for the average user. It’s a professional tool for highly complex work.
I would compare it to a extremely high end $15k PC, or an expensive pro-grade video camera, or a freight train, or a …
I would say at least 95% of the global population will not encounter a situation once in their life where it would be actually useful/warranted.
WinstonSmith84 1 days ago [-]
I wish Fable were as good as you make it sound. A plan created by Fable is good, but in my case, it always contain issues caught only when it's reviewed again (whether by itself, Opus, Sol etc.). That's (almost) not different from plans created by Sol, GLM 5.3 etc. The one thing where it's genuinely better is the front-end, but then again it's far from perfect, it just needs less iterations.
dwaltrip 24 hours ago [-]
Fable does far better at considering the whole picture, weighing options, and making suggestions during architecture or refactoring discussions.
It’s more reliable and makes less dumb errors than Opus.
It still messes up, of course. But for my working style, I definitely prefer it.
mapontosevenths 23 hours ago [-]
Exactly. Use Fable to draft the plan and make decisions. Less capable models to implement the plan.
That said, Opus 5 is broken. Use 4.8 or another vendor for the build agent.
enraged_camel 1 days ago [-]
>> but in my case, it always contain issues caught only when it's reviewed again
Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.
throwaway63467 1 days ago [-]
It works much better on regular software development e.g. for complex refactoring where cheaper models would produce a lot of garbage results.
bentt 1 days ago [-]
Yeah this is a fair point. I only go to it when I have some big architectural problem I want its help in working out. Or a super nasty bug.
tcp_handshaker 1 days ago [-]
You have a $3 trillion bubble riding on this not being true.
jamiek88 1 days ago [-]
Right? If we’ve already reached ‘good enough’ then there’s rough waters ahead.
Yizahi 1 days ago [-]
I have a sneaking suspicion that someone at Google may be making the same bet, looking at the faster and faster Flash models which provide acceptable results to a lot of people (outside of coding).
deadbabe 1 days ago [-]
For Anthropic. But not for the AI industry at large.
Cheaper, more powerful AI will continue to expand the bubble. Projects will get more ambitious. Everyone will build out their own custom little software. Code diversity expands and requires even more AI.
byzantinegene 14 hours ago [-]
most of the bubble right now is fueled from circular financing of Anthropic and OpenAI
deadbabe 14 hours ago [-]
It’s not “circular financing”, it’s cooperation between companies.
porridgeraisin 1 days ago [-]
The bubble is _not_ on models becoming more intelligent and solving arc-agi-999.
They are already good enough at what they mechanically are.
You have to use the right harness, right verifiers (automatic where possible, human where not), etc much much more specific than a generic one like claude code or codex, and it will also be able to work within constraints and be the "proposer" of an imaginary optimisation problem and an excellent one at that. But you have to frame the task at hand in that manner or maybe even reorganise the task you do itself so it is more amenable to being framed that way. If you use it this way, it is _already_ massively economically useful. But it will take many years for it to actually be usable in that way, since you need DC capacity to come up first which is few years away and also well, massive organisations that have to integrate these will usually take many years to do so.
It is also useful albeit less so in cases like general SWE, where you still need a human in a loop for non-verifiable requirements, and also in other general usecases where information retrieval is too intractable and you need to carefully use LLMs as a component of the overall system.
I am not saying Fable or whatever the biggest models are are useless - they will certainly be useful for tasks at the frontier of the day - which is today complex exploits and open math problems, and well, tomorrow it could be something in biotech. But this is not what the entire bet is on at all - just automating day to day drudge at the tens of thousands of massive companies and governments we all know and love is more than enough. With the right training data (which _also_ is a bottleneck and takes time) you could even automate certain processes entirely. Sure, if we get a crazy medical innovation and end up saving trillions in healthcare great, but that's just a bonus.
None of this is to say that I think there is zero sketchy financial engineering going on
krupan 1 days ago [-]
The bubble is based on the promise that these LLMs will cure cancer and find the solution to global warming. The pragmatic users of these tools (like you seem to be) are enjoying the subsidized use of the tools right now, but it's not a sustainable business model
bonesss 16 hours ago [-]
Sam Altman says a lot of crazy stuff, but I don’t think that is where the ‘bubble’ is coming from.
In recent months we’ve had the first automated unmanned amphibious assault, the daily drone count in our hot wars is jumping by leaps and bounds, and arms suppliers are promising future drone shipments in the hundred thousand unit range. The ten year picture for reactive combined swarm intelligence on the battlefield is promising to be widespread, highly lucrative, and in need of constant adaptation to near-peer efforts. Datacenters in space are dumb, datacenters in space to power orbital weapons networks and rapid response capabilities make sense.
On top of that we have international trade, scalable customer service, and a first pass 80/20 answer for businesses focused elsewhere. Shitty, maybe, overpriced, maybe, but useful enough our grandkids are gonna use ‘em.
In both cases, as well as potential new LLM-like tech, there’s an argument to be made for being a leader now to dominate the future. That means compute and tech positioning, and memory & GPU deals.
YouTube was a money loser, Google was ‘losing’ money on them for years, YouTube didn’t have a sustainable business model. YouTube was the biggest, though, and whatever premium Google paid to be #1 then meant they were #1 when the online video business model matured. Now they’re printing money with a platform outcompeting news, social, and video platforms.
porridgeraisin 15 hours ago [-]
Sam altman himself has sort of cooled that rhetoric down - see his recent podcast (frankly I forgot who it was with, I only saw clips of it - but he was wearing orange-brown and had an applovin mug). Maybe it is to contrast against anthropic where they are leaning hard into the EA party line, not sure. This guy always seems to have an ulterior motive behind everything though.
One line I remember was that he said the disruption is not happening the way they thought after GPT4 due to inertia bla bla and that it will be slower gradual change and they got the timelines wrong.
And yes on the defense usecase. That is the main reason governments are giving a hoot about AI. Orbital datacenters too, I know people working on the Indian one, it's entirely for defense usecases. Basically for missile stuff.
fluidcruft 23 hours ago [-]
Are you sure the bubble isn't riding on the idea that satisfying the electricity demands will bring fusion to the world?
anon373839 24 hours ago [-]
The issue is that, if there is a “good enough” point approximately here, it is only a matter of time before models become small and efficient enough not to need all those data centers. Though, it should be good for companies that sell computers (like Apple) rather than putting a toll booth in front of a pile of numbers.
vatsachak 1 days ago [-]
I hate to be brusque but this is cope. GPT 5.6 Sol is just as good and cheaper
nullbio 12 hours ago [-]
1000% cope.
rowanG077 24 hours ago [-]
I think highly complex work and “professional” work are basically completely orthogonal. You can have highly complex work you do as an amateur, where AI can be very useful. For example working through a difficult mathematical problem, building or contributing to an operating systems, or researching a highly technical topic for a hobby project such as microscopy, chip design or lithography.
And you can have extremely simple work that you nevertheless have to do as a professional. For example, drafting a routine customer email, summarizing a meeting, formatting a report, filling in standard documentation, or making a trivial code change.
So I don't think the $15k workstation / professional camera analogy really holds. Those are specialized tools whose capabilities are mostly useful within a fairly narrow domain. A general-purpose AI can be useful across thousands of completely unrelated tasks, including one-off problems encountered by ordinary people.
a11r 20 hours ago [-]
There is a difference between needing frontier capability because one is solving a truly open ended problem, and needing a reliable workhorse model to do something well understood. Local models (like Qwen 3.8 27B) have gotten so good that they can do all routine tasks at a fraction of the cost of frontier models.
pmdr 16 hours ago [-]
That model costing $3/M output tokens on Openrouter is a mystery to me.
nicce 15 hours ago [-]
32GB GPUs are getting expensive so that also affects the price.
pmdr 14 hours ago [-]
I didn't think about GPU size. So they don't run this small model on B200s or something?
nicce 9 hours ago [-]
This is open model intended for local usage. So, it doesn't matter. If the end user wants to run this on their own hardware, cost is defined by the electricity price + the price of the processing power. Providers can increase the price as long as the users don't switch for buying the hardware themselves instead.
lbriner 7 hours ago [-]
It's a bit of a non-headline. We are a period of massive up-take from people who aren't really familiar with what works well for their use-case and how much that costs.
I have tried fable, gone back to sonnet, tried to use agent mode to mix and match and various other combinations. If I decide that Fable is the only one that can save me hours or days of time, I will go back to it and pay the cost. If the cheaper or free ones do well enough, I will stick with them.
When you see the price of hardware needed for these models, we are generally not paying very much towards that price so I expect prices to start ramping up as the honeymoon ends and these companies' investors get itchy for their ROI.
mirekrusin 22 hours ago [-]
Poor analysis.
Fable is not used because it has extra retention requirement that corps can't sign off so it stays disabled for everybody in many cases.
It's also not winning on day to day work against Opus 5, which is simply available as there are no extra retention requirements and no extra paper work to do with legal.
Devs also don't like that Fable refuses to work on half of their prompts.
Then comes better pricing and on-par capabilities from competition.
TylerE 21 hours ago [-]
> It's also not winning on day to day work against Opus 5
In what universe? Hell, Opus 5 seems to be worse than Opus 4.6 at almost everything.
mirekrusin 16 hours ago [-]
Sure.
lnenad 1 days ago [-]
As a small background, I have a local server and I've been trying out different models with different inference engines, quants, configurations etc... I'm also using Opus and Sol at work consistently. I've used AI since the first wave, first as a toy, then as a highly specific tool, last 6+ months as the primary LoC generator.
This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...
When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use Fable still *feels* much better. I think benchmaxxing the new open weights models is ruining the value of benchmarks, if they ever had any. When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. And considering the point I'm poorly trying to make, that Opus doesn't feel like frontier anymore, this is probably the last month of my Claude subscription.
philipbjorge 22 hours ago [-]
> When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
I've wondered if this is part of why we don't see the reasoning traces for Anthropic's models before -- Open models might just be accurately surfacing how the sausage is made.
lnenad 17 hours ago [-]
I'm assuming it's definitely part of the equation, but considering that I'm getting more tps but still waiting a lot more time for code to come out I'd assume it's not a 1:1 comparison. Plus I'm running quants, maybe with full precision it's better.
jeffnash 1 days ago [-]
Hasn't it always been the premise that intelligence would get cheaper? To me, on the enterprise side, it seems like firms are finally getting the memo that, whether you are locked into the Ant/OAI ecosystem or not, you don't need the smartest, most expensive model to do every single task. This is a good thing for overall adoption. Whether that trickles down into regular user behavior, especially with subscription pricing, remains to be seen; even though I intellectually know I don't need Sol for a simple refactor, I am sometimes hesitant to choose Luna/Terra, as it's hard to accept using something positioned, even implicitly, as 'worse'. Remembering that the smaller models tend to be faster is what usually pushes me over the edge.
Anthropic in particular is much more compute-constrained than OpenAI and SpaceXAI and has relied on partnerships to provide inference. This reality factors into their pricing and usage limits (they started 'adjusting' the 5-hour limits during peak hours, and it certainly wasn't an upward adjustment). Accordingly, this is presumably what Anthropic wants, given they develop and release the lower-end models, suggest users use them in various nudges within their product, position the bigger/more expensive models as "For the most complex tasks" in their UIs, and so on.
margorczynski 1 days ago [-]
> whether you are locked into the Ant/OAI ecosystem or not
I think the problem (for Ant/OAI) is that there is no sensible lockin or moat. LLMs are essentially interchangeable and stuff like a harness doesn't offer enough value on its own for someone to be locked into using one of them.
Now with the onslaught of the Chinese models that offer almost the same quality for much less money they have a very serious problem on how to proceed. Investors now might be looking through rose tinted glasses but their patience has its limits.
jeffnash 24 hours ago [-]
Agreed 100% for the consumer case: an empty chatbox is just about the least sticky surface I could ever imagine. I saw a mobile interstitial ad for Kimi recently whose hook was basically "Tired of paying for expensive ChatGPT? Download the Kimi app, it's the same thing but cheaper". I myself bounce between token subscriptions like no one's business and use Pi/OMP for maximum model flexibility when coding (and it's a few env variables or lines of (TO|YA)ML|JSON to switch providers in Codex, Grok Build, CC). I even self-host and try to use OpenWebUI + CLIProxyAPI when I can for all my chats.
Enterprise is a whole different ballgame IMO with countless technical, compliance, and employee adoption considerations that add friction to switching. It's also where both Anthropic and (as of last week) OpenAI get the bulk of their revenue, and, incidentally, the venue where US Government regulations on Chinese models would have the most impact.
kollegekid 17 hours ago [-]
This article totally misses the fact that fable is not ZDR!!! no serious large corporation can use it (or at least without a lengthy legal review)
wiredbox 10 hours ago [-]
Fable stuggles to attract users because it's so bloody risk averse it's ridiculous. At a hint of something that it might interpret as a red flag (e.g. cyber) it will downgrade straight to Opus. As someone who's utilizing LLMs namely in the context of infosec, Fable has simply been unusable.
nullbio 12 hours ago [-]
I've found myself using Kimi K3 API for things that ChatGPT is not good at (primarily AI research & development, because I swear they nerf their models for this - along with Anthropic), and I see absolutely zero reason to ever use Anthropic's APIs. The only reason I can tell they still have any form of momentum is because of sunk cost fallacy from the users who still use it.
czhu12 8 hours ago [-]
Its like the world speedran the problem with phones where at one point, annual upgrades really didn't make a difference between everything was so good already.
At this point, every model writes code about as good as I'll need for the stuff that I'm working on, and with the right harness, I'm perfectly happy to let them cook, until it comes up with something working. Whether it takes 10 minutes or an hour really makes no difference to me.
mpweiher 14 hours ago [-]
This looks like strong evidence against the claim that local models will never be competitive, because hosted frontier lab models will always be better.
The frontier models may be better, but who cares if the last generation of models are plenty good enough for what you want to actually do?
And the best local models are quite competitive with those older lab models.
chermi 6 hours ago [-]
If they just told opus to stfu a large fraction of people wouldn't even consider leaving
jatins 18 hours ago [-]
For _most_ day to day knowledge work (writing, excel, filling forms, coding) current SOTA models are good enough. There are diminishing marginal returns from paying more in my opinion.
If you are disproving Jacobian Conjecture it makes sense to be on SOTA, but for writing Golang and Typescript, faster sol/fable/opus class models are imo more likely to get user interest than the latest frontier.
nrmitchi 22 hours ago [-]
Anthropic's issues, as I see them, are:
1. They acted as if they were so far ahead capability wise that they could stop listening to their users.
That is basically it. It is an extremely common belief that Opus 5 acts like a condescending wannabe-thought-leader, yet Anthropic's response to this is largely been "You're using it wrong. Try deleting all your config files".
They released a "concise" output format, but pretty much swept it under the rug, despite it being one of the loudest complaints about their models. It also, generally, does not work as described.
Not only did Opus 5 get much more difficult to work with, but it got, from a customer view, significantly slower. The token rate might be the same, but if every interaction takes 50% more tokens, it's 50% slower.
All this time OpenAI has released a slew of new models, lowered the price on them (which, bluntly, 80% of user count don't actually care about since they're on subscriptions), increased subscription capacity, and increased response speed.
It's not about cost. It's about Anthropic being the frontier-lab version of the marathon runner who decides to celebrate to early, and then loses the race.
nik736 11 hours ago [-]
For me Opus 5 was the nail in the coffin. Fable without the restrictions was a great model, but became unusable with the security guardrails. Opus 5 became so bad and slow it's unbearable. And since Fable falls back to Opus all the time it was time to switch. Me and my friends are calling it Slowpus by now... Up until some weeks ago Anthropic was the king, but we all switched to Grok 4.6. It became slow as well but is not down all the time, is a solid model and paired with another reviewer model it's a great daily driver.
Rover222 8 hours ago [-]
Grok 4.6 is great
cheeze 6 hours ago [-]
Similarly to the cars, I won't use it on moral grounds.
Rover222 4 hours ago [-]
okay...
sajithdilshan 1 days ago [-]
Every software engineer in my company uses Claude code heavily. However we’ve never enabled Fable and only use Opus, Sonnet and Haiku.
Nobody has complained and seems like for every use case we have Opus is more than powerful enough, especially with Opus 5
preommr 24 hours ago [-]
> Every software engineer in my company uses Claude code heavily.
I find it funny how OpenAI got caught lacking for a very brief window, but it turned out to be a very critical turning point.
Like a guy that that's at the top of their game the entire year, and the one day they have the flu, the CEO does a surprise performance review.
sajithdilshan 24 hours ago [-]
Yes, the critical point was end of last year, beginning of this year. Especially around the time Opus was released.
dominotw 20 hours ago [-]
i think they were trying to play a different game.
dham 24 hours ago [-]
A model is just another thing to plugin to a harness. I don't give it much more thought than that. If developers are still caught up on Claude Code, or Codex that's just not a long term thing. It's best to develop workflows locally and in the cloud with open harnesses. I know this will be the future because that's how it worked on every other system that developers use.
Sure there are Microslop and Oracle db users but most of the world we live in is Postgres and Linux. That's why I think most companies will run llm's like that.
sajithdilshan 24 hours ago [-]
That’s not true. Before AI, I have been using Jetbrains IDEs as far as I can remember. Also have been using MacBooks for work since my first job. You don’t have to generalise everything. If a particular specialised tool is good at its job just use it instead of re-inventing the wheel
dham 23 hours ago [-]
Models will be a commodity
> good at its job just use it instead of re-inventing the wheel
Exactly why is everyone reinventing a harness every month. There will be Microslop / Oracle harnesses and there will be 1 or 2 open source ones that win.
cromka 16 hours ago [-]
Paying 100 USD a month and getting "Claude is at capacity now" when you need it for work doesn't help.
You can't make it an indispensable dev tool if devs cannot rely on it. They will absolutely jump the service to s more reliable one and that was Codex for me.
P.S. Claude is indeed down right now for me.
tosh 16 hours ago [-]
The chart ends in July, would be interesting how model adoption looks like in August
Also Sonnet 5 was released June 30th, seems to be grouped in with 'other'?
Opus 5 is by far the best model for everything that matters to me. But it’s just too expensive. Terra 5.6 is bearable for everyday tasks, so now it’s my default.
atleastoptimal 22 hours ago [-]
Like 90% of consumer AI use is stuff like "please write a summary of this PDF for me", or "Find a cheaper version of this product online", eventually the returns on model intelligence taper off for these kinds of tasks. However for the kind of frontier tasks they are testing the models on now, like math, science, etc. the marginal returns on increased intelligence are huge.
k8si 7 hours ago [-]
Our org can't use Fable bc they require 7-day data retention to be turned on and my company won't do that. So it might have something to do with that rather than actual lack of demand.
exabrial 1 days ago [-]
Perhaps nerfing the cap out of your best model for press attention and hosting valuable features like thought traces isn’t such a great business model?
jstummbillig 13 hours ago [-]
The reading seems a little presumptuous to me. We do not know what Anthropic expected. This is mostly a matter of performance/cost and they clearly focused on building the Rolls Royce. How much adoption do you expect, when you build the Ferrari? And how many Rolls Royce can you even deliver?
Saying "people are not buying that many Rolls Royce, instead they buy a lot of normal cars" is kind of duh. Anthropic is clearly still operating at inference capacity.
stephencoyner 8 hours ago [-]
I think it’s clear that most of their revenue is enterprise and enterprise demands zero data retention, which you can’t have with Fable currently. Non starter. This isn’t a price thing
hmokiguess 24 hours ago [-]
I guess soon they will see the full picture
dominotw 20 hours ago [-]
the picture is now clear
motbus3 16 hours ago [-]
My first reaction to Fable was:
"This is good for the bang"
But as time passed the brittle software I noticed that without strict guidance it builds poor and brittle software.
If I need to write all the specification so it follows it, I might just write the code or use a cheaper model.
Opus 5 and Fable 5 have been quite disappointing.
I used sol, terra and Luna and I think they and the first two are good for first code reviews. Luna is not much better than deepseek V4 flash
zkmon 12 hours ago [-]
There is hardly any invention-based moat a startup can have these days, in this fully connected world. The only moat is laziness of people sticking to known products and solutions, real world assets that others can't replicate easily and ability to do some real world work.
mrwh 22 hours ago [-]
Are we at the stage yet where a super-strong model that needs oodles of safety protections to stop it outright hacking you is strictly worse for day-to-day tasks than a much cheaper model that's simply not competent enough to be dangerous?
sreekanth850 19 hours ago [-]
It all started when OpenAI began focusing on Codex and coding while sunsetting Sora, which helped free up a lot of compute and resources. 5.6 was the final nail in the coffin.
apparent 22 hours ago [-]
I have a $20/mo subscription and have found myself getting limited every few minutes of late. I'm just building a simple website, but after about 15-20 mins of back-and-forth, it tells me I'm tapped out for another 5 hours.
This wouldn't be so bad if it didn't make mistakes periodically, especially when it's about to tap out. I get the sense that if I upgraded to the $200 subscription it would get me a lot more usage, but it would still run into these issues anytime I sat down to work for a few hours.
I'm just using medium effort, so it's not like I'm on high all the time.
Aeolun 16 hours ago [-]
More recently, after building something with direct chatgpt access I’m just baffled by the difference in speed between sol and opus. Opus hadn’t even finished making a plan, and sol was already done with executing. It’s like…
Then on top of that opus slurps tokens like Anthropic is afraid they’re losing money. Read 200 lines of file, +20k tokens. Excuse me?
fbrncci 20 hours ago [-]
I have yet to even try Fable or Opus 5. Just looking at the hype they put out prior to the release, then the whole way of releasing these models as well as the pricing just puts me off to get used to it and then needing it. And this while so many good models came out without any hype, botched releases and significantly cheaper. Anthropic really shot themselves in the foot. I went from using sonnet and opus models for 50-60% of my daily token usage to 10-20%
Metacelsus 21 hours ago [-]
Well, as a biologist, Fable is still completely unusable
colingauvin 19 hours ago [-]
I can't even ask it how to make toast without zeroing out all its memories of me.
claudes_bussy 22 hours ago [-]
I have 200MAX subscription for Anthropic and a RTX 5070TI 16gb GPU. I only use Claude.AI web chat and I only use Opus4.7. I build my prompts on prem with Qwen3.8:latest and copy paste them in claude.ai. I never type anything into claude.ai I am merely a copy pasting monkey. If I need to adjust something I do it from the the on prem prompt. I download code bundles from Claude.AI and push them to my git repo I host. I instruct claude to not do any testing and only let me do testing on my machine so i am not wasting claude resources. The point is to minimize the claude agents from doing anything but specifically doing coding. I can do this for 12 hours a day and reach about 80% of my weekly usage. Once I get to 3 hours left on my weekly I start a fresh chat session with Fable5 and have it do blind code reviews on everything. I only do this because I want to max out my weekly allotment and fable will get me to 20% in 3 hours on a 200MAX sub. Sometimes Fable5 does something interesting but for the most part my on prem static and dynamic code analysis have kept everything buttoned up.
mikert89 23 hours ago [-]
im using fable almost exclusively, i just buy a new 200$ license if i run out of capacity. its so much better, easily worth it given how much work I get done. I built a code, deploy, e2e test loop with my full aws infra, fable just implements linear tasks constantly, 5 at a time, tests the whole thing end to end
if people dont see why they need a model this smart, they probably arent using ai enough
HDBaseT 23 hours ago [-]
People have realized you don't always need a Fable level model. Majority of my work is sufficient with ChatGPT Luna which has effectively unlimited usage on the $100 ChatGPT plan.
I am not doing awfully complex tasks though. I imagine a lot of other people are in a similar boat, either switching from Claude to ChatGPT or even just min-maxing DeepSeek V4 Flash 0731 or similar.
mikert89 23 hours ago [-]
the code quality is higher though, so the value compounds. i dont need to watch as closely to what its doing, so its more autonomous
apt-apt-apt-apt 21 hours ago [-]
Isn't there a risk of getting banned if you do this (multiple accounts to bypass usage limits)?
mikert89 21 hours ago [-]
Idk I never worry about it
STELLANOVA 17 hours ago [-]
We never got updated Haiku model and it's a shame. Not everyone and everything needs PHD level knowledge/reasoning, in fact it's really rare you need Fable level of knowledge/reasoning for vast majority of users...
piker 17 hours ago [-]
It does seem like diminishing (perceptual|valuable) returns to intelligence is antithetical to exponential capitalization growth. Even if we can get the models to Einstein level, maybe we just don’t need that to, say, mow the lawn.
throwaway63467 1 days ago [-]
Yeah I mean if I run out of tokens every couple of hours and have to pause my work or shell out more money I’ll switch to other tools that don’t have this problem. Though they turned this down a bit it seems, I can work with Fable reasonably now and I enjoy it actually. I think they were just testing out how much they can raise the cost without users leaving when having the best model. I guess not much after all!
bastawhiz 8 hours ago [-]
This is maybe an unpopular take, but I don't think this matters for Anthropic. I don't think their immediate goal is to get users onto their biggest and best models. Sonnet is more than enough for many average users and their use cases.
At this point, Fable is really just an experimental model (as it should be). It can do very useful things, but is it a broadly good general purpose model? Definitely no. Most users don't have problems hard enough for Fable outside of coding very large and complicated projects. Most users don't have 45 minutes to accomplish a task Sonnet can do well enough in five. There's not a PowerPoint in the world where Fable is the right tool to build it.
The secret sauce is going to be in letting Sonnet decide to delegate to Opus and Fable when they're the right tools for the job. But you can't train a model to do that until the bigger/better models exist and you can observe how your users actually take advantage of them. I'd bet money that's exactly what the rlhf going on at Anthropic looks like right now.
Economically, it makes sense. Sure, on paper you want users burning as many tokens as you can. But pushing users into burning tokens and taking a long time and getting a meh result is far worse then giving them the "fast and good enough" solution that occasionally burns more tokens automatically when the problem requires it, and getting a higher quality result out when you do. From an infrastructure capacity perspective, this is the dream: you stop measuring cost [for Anthropic] per token and measure cost per outcome, allowing you to use less hardware to accomplish the same abstract units of work.
t0bia_s 16 hours ago [-]
Claude Opus 5 did plugin for Photoshop for $4.5 in three responses.
DeepSeek v4 Pro did same for $1.7 in 7 responses (peak-off times).
For doing much complex task, I would be super nervous about using Anthropics models.
sagex 22 hours ago [-]
With open models delivering close to frontier level capability, I really don't see how these labs can only focus on having best model to sustain their business. Although these model's usage will grow, I think and hope that the serving will be distributed among many players.
aurareturn 20 hours ago [-]
On subscription plan. I only use Fable. It’s been my exclusive coding model since its release. I refuse to use anything else for coding.
For AI agents, we have been mostly using GPT 5.6 Terra.
jbverschoor 15 hours ago [-]
It’s slow. Unbearably slow. Back to using a mixture of other tools.
Did they not learn? Performance is a feature
tschellenbach 23 hours ago [-]
I see variations of this post all over X. Fable doesn't get much usage, since Opus 5 is like 99% as good at a lower price point. And perhaps more importantly also faster.
asimpleusecase 17 hours ago [-]
Open router is the answer, get the intelligence you need and have the flexibility to not get trapped
anonu 20 hours ago [-]
so whats the best and most efficient coding harness and against which model? what are folks doing to keep costs low? I spent $1000 just this weekend on my personal projects for sota Claude but I feel like I can probably get much more juice if I start looking elsewhere.
and resources or tech stack tips from HN?
ThunderSizzle 12 hours ago [-]
Buy a card. Run your own. It'll be some time to optimize it, but an AMD R9700 is one of the cheapest by $/vram. Sadly, its been increasing in price slowly, and there seems to be some inventory issues, but I've switched to it exclusively. I hope to get a 2nd card to run other workloads simultaneously.
I've been running Qwen27B Q6 with 130k context at 30tps. It's not bad.
I think with an "autopilot" mode in pi, and sub agents to handle context window management better, I'd envision you could get close to unattended workflows. I haven't quite gotten that far though.
Negative is you have to do it all, and the temptation to tinker is real.
mupuff1234 23 hours ago [-]
Which is why they are rushing to an IPO.
cmiles8 22 hours ago [-]
The foundation model companies can’t survive a token price war. If that’s where we’re heading get out your popcorn.
mattwad 20 hours ago [-]
Claude will refuse to visit sites with robots.txt, make a graphic spoof of an iOS game, or even fill out an employee survey on my behalf. ChatGPT is always happy to oblige, no questions asked.
I can respect the guardrails - I also can see why OpenAI may not have much control over their models - but I need an AI who will do whatever I ask and not play judge and jury.
luciana1u 17 hours ago [-]
the expensive model is losing to the cheap one because most people's problems were never that hard. the expensive one is for the problems you only find out you have after the cheap one works.
firemelt 8 hours ago [-]
i hope they are collapse
SomeHacker44 21 hours ago [-]
I am getting close to dumping Anthropic. I like their models in general, but boy I have hit their "F you, I ain't gonna help" too many times now on innocent things. Ain't nobody wanna deal with that. I have never gotten that from Antigravity and if I did I woupd tey Codex then go to Openrouter and leave US models behind.
I mean, who says "screw you" to requests to get 35+ year old vintage computers working? Claude, that is who. Its guard rails are so stupid. I hear people trying to do simple mailing list management hit it too.
I am just about done with them.
jaykru 18 hours ago [-]
GLM 5.3 will probably do your vintage computing work without fuss.
epsteingpt 18 hours ago [-]
The rate of model improvement has slowed, and may not recover.
It's unclear if Mythos2 or 3 or whatever they're calling their next model will be an improvement for most common enterprise use cases.
LLMs can't solve basic things (writing non-slop documents, understanding context without massive handholding) and for coding other models are quickly becoming 'good enough' without the same cost and nannying.
That's why Anthropic is 'stealing' workflows.
But it turns out it's much harder to push adoption when your users don't really want to use your product.
Code was a unique use case where the code luddites were loud but a minority - most people don't want to update 300 cases of variables across their code base for a name change. Most don't want to write unit tests.
There are a few use cases where that will happen (law is next, maybe quant finance) - but otherwise most companies are throwing money into a pit and getting 0 return.
It's a very interesting race and state of affairs, but Kimi K3 and likely the next DeepSeek models will put the high price token affair to rest.
Unless of course, mythos / next model really does solve some universally applicable problem that people want it it to do.
qwerty2020 7 hours ago [-]
Not only is it incredibly expensive for the intelligence provided, it's slow, overly verbose, and outright refuses its users all the time. The level of general hassle with Anthropic nowadays is not worth it...not to mention Claude's infuriating dialect.
ranang 16 hours ago [-]
Until now, I've been paying Anthropic for the $200/month plan for over a year. This morning I was once again hit with "You've hit your monthly spend limit · your weekly limit resets 8pm […]". I am so frustrated and tired of being treated like a lab-rat to see how much I am willing to pay for less and less LLM access.
I've now bought the $200/month plan from OpenAI and simply ran `/status` in Claude to get my session ID, then asked Codex in the same directory to "Please take over the work started by Claude Code with session ID `[…]`." This seems to work like a charm. It also seems that the Codex weekly and monthly budgets are more generous?
Unless Anthropic adjusts their customer-abusing behavior I think I will end my subscription with them soon.
isolay 20 hours ago [-]
What a shame, they cooked up all their marketing lies in vain.
x3n0ph3n3 23 hours ago [-]
Anthropic's privacy deviation (ZDR) for Fable is why my organization forbids usage of Fable.
latentsea 21 hours ago [-]
Qwen3.8 is all you need.
Frannky 19 hours ago [-]
Omp + 0x alpha is free via openrouter and opencode go
surume 12 hours ago [-]
The reason I try not to pay for Anthropic is because half the engineering questions I ask get flagged as "too dangerous". Examples: spraying liquids at relatively low pressure, creating small AI models to detect objects via Raspberry Pi cameras, and calculating collisions between moving machine parts. Claude blocks me on almost EVERYTHING, even though the questions are completely valid science and engineering questions. I see NO REASON to pay for Claude when Kimi or GLM's quality it almost as good and I don't get rejected all the time for absolute nonsense. Anthropic is trying to be the uber-safe children's chemistry set where its ok if a kid drinks all of the chemicals in it while playing with it. That's not how you run a successful business. This is just the natural result.
thelastgallon 22 hours ago [-]
People understand the law of diminishing marginal returns.
j45 7 hours ago [-]
It's hard to trust if cloud providers will be available any given day, and the new angle of sanctions for new models seems to have people exploring alternatives.
I'd use it if I could get through a 5 hour session without exhaustion my usage limit.
hncbw02z5a 11 hours ago [-]
Every word of this
22 hours ago [-]
Rover222 8 hours ago [-]
My work pays for unlimited plans on whatever models we want, and GPT 5.6 sol has been the workhorse for weeks. Fable is still great, of course, but it's slow, wordy... I don't know. Sol is feeling like the leader at the moment. (coding web dev)
vonneumannstan 8 hours ago [-]
Focusing on consumers is the wrong idea. They're a B2B company.
segmondy 21 hours ago [-]
Give me a break, Anthropic is expensive and their CEO is not the nicest guy around, and the games they play with other people's money/their API is not fair. I stopped using Anthropic and OpenAI once they started calling for regulation of open models. I have survived locally since LLama3-70 days and have been surviving fine. If I was to pay for cloud models, it definitely will not be Anthropic, Fable or not. From what I have read, the best AI model will not even comply with requests most of the time because it or/and Anthropic supposedly knows what's better and safe for you.
t1234s 21 hours ago [-]
Do neoclouds win out in this scenario?
dboreham 9 hours ago [-]
Since this seems to be a thread for people to criticize Anthropic let me add my $0.02 that I'm a very happy customer and see none of the various terrible things everyone is complaining about. Except Fable refused to give me some flags for nmap. That was pretty annoying.
RayVR 17 hours ago [-]
I’m constantly disappointed by Anthropic’s models.
I become more disillusioned every day.
It seems like, by having the models write Python code, they tend to write Python code like an average developer. Which is to say, quite bad.
Add in the complete failure of the models to adhere to instructions in Claude.md, memory files, and added multiple times in prompts, I’m wasting huge amounts of time fixing bad design decisions that the model just slips in.
1saadcodes 21 hours ago [-]
Honestly, this makes a lot of sense to me. Devs that are good at their job don't need the absolute best model for every task, and if a cheaper model gets the job done 95% as well, it's a pretty easy choice
sebastianconcpt 23 hours ago [-]
The market correcting itself.
708145_ 16 hours ago [-]
All of the version 5s have been truly disappointing.
- Fable, nerfed or whatever, too expensive and not fully included in subscriptions.
- Opus, neurotic (excessive) slop machine.
- Sonnet, way too token hungry, cost much more than 4 series.
Their almost daily outages does improve the experience. Anthropic really have messed up this year.
vikramkr 14 hours ago [-]
"spending in fable 5 has plateaued"
Bro if you want us to spend more on fable let us use our whole damn rate limit for it. Enterprise is a different ballgame obviously but for the subsidized Claude code users they literally cap it
cedws 13 hours ago [-]
Huh? Fable is not benching as the "best" model anymore, that's probably why it has low usage. Opus 5 is supposedly the best one now, no?
emsign 15 hours ago [-]
When are the hyperscalers finally collapsing? I can't wait! I hate their hardware hogging and land grabbing so much. AI has to be local it makes no sense otherwise.
LoganDark 16 hours ago [-]
It seems the people who aren't attracted to frontier models are the people who want to do the same they've always done, but just with new tools. I want to see what new things I can do, so of course I always want the latest and greatest. The stuff I've always done, I've already beaten to death.
alpaca9 11 hours ago [-]
'People prefer models that are as good or better at a lower price, what a surprise!'
diogenescynic 16 hours ago [-]
Markets been waiting for something to sell off or correct in the short term over. Looks like it may have found its justification.
diogenescynic 8 hours ago [-]
Guess not! Markets shrugged it off.
Mistletoe 18 hours ago [-]
If you are still investing in these companies or plan to in the IPO, the financial ruin you experience is your own doing, you are ignoring every sign that this isn’t going to work out. None of these numbers make sense and point to AI being a commodity with razor thin margins and a race to the bottom. Would you invest heavily in a toilet paper company that took massive amounts of power to produce each version that is 1% softer or stronger every six months?
19 hours ago [-]
guluarte 9 hours ago [-]
nobody in my company wants to use opus 5 and fable is too expensive
dejan_ 16 hours ago [-]
When I see a paywall it's always some scam. Always, both the article and the LLM.
rvz 21 hours ago [-]
People here won't believe this, but Anthropic will begin to decline after their IPO when everyone runs to good enough cheaper models to save on token spend.
moralestapia 22 hours ago [-]
Not just that but also Opus 5 is real trash. It's free for me (company pays) and I still prefer to use other models.
This article is a tiny bit deceptive -- it should read "struggles to attract API users". I imagine that subscription demand for Fable 5 is very high. But clearly Fable 5 is a much less attractive model for enterprise service use; you can clearly thank the Trump administration for some adoption hesitancy as continuity of service for Fable has been demonstrably insecure. We do not use Fable at the service level at our organization for several reasons: cost, service continuity, and data privacy. But we use Fable extensively for code development.
lenerdenator 7 hours ago [-]
If I were an entrepreneur, I'd focus on getting a platform up that runs open Western models with good management and governance tools, and a decent coding agent front-end.
The whole "we're going to replace all of your workers with an agent while also making AGI happen" thing was a bad idea in the first place; one that could not have gotten anywhere outside of Silicon Valley. People just want tools that they can deploy at scale, not to be your beta testers for the singularity.
TechSquidTV 22 hours ago [-]
I've been saying since the beginning. Models. are. WORTHLESS. If your company depends on having the best model, you have lost.
verdverm 1 days ago [-]
I've always wondered why everyone flocks to SV's latest darling company. Have we not learned from our history of glorifying these SV darlings that turn hostile?
Eufrat 1 days ago [-]
I think the glib answer is, “Greed blinds all”.
Most of the people pushing this are just hoping that they can cash out before the hype pops and financial gravity crashes the party. Sam Altman recently claiming that the singularity is here is so stupid on its face he should just be treated as what he is, a huckster.
None of this stuff ever made any sense on what it was being sold initially. It was always insulting that the media and business leaders tried to argue that the tech could replace entire call centers or vast swaths of entire industries.
People keep arguing, but it will or it has based on extrapolating certain, reasonable use cases. Klarna has shut up about replacing call centers with bots because Markov chains with memory only can do so much.
brookst 24 hours ago [-]
Some of us want to get work done and don’t feel the need to either glorify or hypothesize about what might happen.
I’m happy with Claude. If they become (bigger) jerks, I’ll switch to something else. I don’t ha e the energy to praise Anthropic today and I won’t have the energy to demonize them tomorrow. The emotional investment people have for/against these companies feels like celebrating or being offended by the weather.
tyleo 1 days ago [-]
Are people flocking to them? I see people buy the products but if you ask I think they are just about as hated in the big techs.
dofm 24 hours ago [-]
A mixture of opportunistic edge-seeking, FUD, FOMO, novelty-seeking, the need to impress shareholders, the tendency of salespeople to believe other salespeople are telling the truth, the ever-present need to stay in front of relentless commodification, and pragmatic curiosity.
TBH I don't think any of that is unique to the IT industry's relationship with the Valley. Other technology-driven industries have a similar worship-ish relationship with a few rarified businesses. But the culture of the IPO exit accelerates all the most short-term motivations to do anything.
Aeroi 23 hours ago [-]
this article reads like a leak from a bank that didn't court the IPO.
Anthropic is the fastest-growing software company in history. They have no issues "attracting users"
ThundeChile 14 hours ago [-]
paywall.
bellowsgulch 1 days ago [-]
Reads like: brilliant Carnegie Mellon University computer science grad struggles to find job where he is not replaced by cheap, inferior Indian labor that still gets the job done, even if it takes marginally longer.
No shit we're all paying 清冲 Flash to do the grunt work. Turns out though, paying 清冲 Flash a few more cents does exactly what Ivy Wasp Pro Mythical does. Crazy how that works.
phyrex 12 hours ago [-]
清冲 is "Qing chong", pronounced "Ching Chong". It's not a real word, it's just racist.
1 days ago [-]
websimapi 1 days ago [-]
[dead]
luciana1u 7 hours ago [-]
[dead]
runtime_lens 14 hours ago [-]
[dead]
coachdaniel2026 16 hours ago [-]
[flagged]
throwaway613746 21 hours ago [-]
[dead]
effnorwood 24 hours ago [-]
[dead]
Taikhoom10 19 hours ago [-]
[dead]
marsven_422 17 hours ago [-]
[dead]
floki165 16 hours ago [-]
[dead]
iandanforth 7 hours ago [-]
tldr: "Company's product is overpriced, market reacts normally."
CurbStomper 22 hours ago [-]
[dead]
Altaba 19 hours ago [-]
[dead]
felixgallo 1 days ago [-]
[flagged]
colingauvin 1 days ago [-]
I am because Claude has figured out I am a biochemist and therefore even asking Fable what the weather is gets me bumped back to Opus, sometimes even Opus 4.8 instead of 5. Kimi? GLM 5.3? DeepSeek? No such problem.
I can literally open a new chat with just "Hello" and it gets bumped.
not_a9 23 hours ago [-]
Memories break Fable 5 for me as well in chatbot. I ask Opus a lot of sec related stuff and now if I even type “hello” in chat it gets insta-downgraded to Opus 5.
jaykru 18 hours ago [-]
Wow, that's totally crazy. Does it happen even if you clear out the "offending" memory entries?
MagicMoonlight 1 days ago [-]
[dead]
felixgallo 1 days ago [-]
[flagged]
dgellow 1 days ago [-]
Fable guardrails are insanely restrictive, are you denying that’s the case?
felixgallo 1 days ago [-]
[flagged]
kami23 1 days ago [-]
Yes that is exactly what's happening.
Claude keeps track of memory if you've turned that on so all conversations are somehow tracked over time Claude randomly will mention that I'm a developer while I'm asking unrelated questions and say oh because you're a developer you might like this or because of my background and infrastructure you might find this interesting and I always get creeped out by it.
They have the concept of incognito chats, but I can see those being worse without context from other conversations happening.
felixgallo 1 days ago [-]
[flagged]
1 days ago [-]
alightsoul 1 days ago [-]
Dario is that you?
colingauvin 1 days ago [-]
It has figured it out through the software that I work on, data that I analyze, questions I ask, etc. I guess I can't prove it's because I'm a biochemist, but that's the only thing that makes any sense. Either way it won't let me use Fable.
dgellow 1 days ago [-]
Im starting to doubt you actually used fable
cmenge 1 days ago [-]
Just cancelled my Claude subscription and used Ox Alpha Free through OpenCode (so that's one).
In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview.
Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro.
It does have issues with availability - but it's on a free promo right now. That also means that I don't know how much it would have cost if I had to pay API prices for it, which might not be cheaper than the subsidized small company / consumer usage, but for large companies paying API prices for either offering, the difference will be considerable.
Some people say it is a new version of Gemini Pro - this is based on some tweets from their employees.
vessenes 1 days ago [-]
Rumored to be mimo
HDBaseT 23 hours ago [-]
Are you sure, most indicators suggest Z.ai with a potentially Flash or Air variant of GLM 5.3?
tyleo 1 days ago [-]
I almost feel like there's some sort of paid campaign going on in Hacker News promoting the Chinese/open models. I feel like every day I'm hearing about how the frontier labs are dead but your experience is the same as mine. My company pays for Claude AND Codex but we never really use the open models for anything critical.
dgellow 1 days ago [-]
Doesn’t matter if your company pays for Claude, Anthropic and OpenAI valuations and expenditures commitment requires them to win the vast majority of the market to make economical sense. And the Chinese competition makes that very unlikely, to say the least. They likely won’t disappear fully but their “free” lunch as the AI darlings is done, on paper. Will be interesting to see how they adapt
ddxv 23 hours ago [-]
Sure, I'm being "paid" by the fact that it costs pennies a day to use Deepseek and if I could afford the hardware then I could run it locally. Meanwhile, Anthropic is obviously a threat to open weight models and actively lobbies the US Government to have them banned or controlled.
So yeah, I and I guess others, are quite active in whatever little way we have available, to up vote new models and share stories.
tokai 22 hours ago [-]
[dead]
felixgallo 1 days ago [-]
[flagged]
hgoel 1 days ago [-]
"Everyone that disagrees with me is astroturfing" is not going to lead you to being in touch with reality.
paulddraper 1 days ago [-]
[flagged]
codexon 21 hours ago [-]
I'm not a biochemist and I have been blocked by fable and opus for "cyber" just for doing things like asking it to ssh into one of my servers, look at a CVE, or do work in assembly. Been rejected to their cyber verification program 3 times already.
openamer 18 hours ago [-]
Great find. Related: we are building OpenAmer, a fully open-source agent that controls the actual desktop (files, browser, terminal via CDP), has persistent memory and A2A multi-agent swarms. Apache 2.0, runs local on Windows: github.com/openamer/openamer
They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling:
"You can only use Fable for a week as a part of your plan" "Be ready! You have to start paying per token!" "Nevermind! we extended it for a couple more weeks" "Wait, now it's up to half your usage" "Ok, now its..."
Most people want to not care. We want our AI like electricity -- Kind of just there no matter how easy/hard is for the supply. You don't want your electricity company to be on the brink of cutting you off any second.
That's Anthropic. You don't feel they want to give you a dependable service for an, albeit premium, price. It's a constant bargaining game. That forces people to look beyond the walled garden. There, they find models that are fine... and without the shenanigans.
It also didn't help that the government yanked it which adds another source of anxiety since OpenAI is on much better terms with the administration and the administration seems corrupt enough that they would mess with Anthropic if they got a big enough donation from OpenAI.
But anyway after Sol entered the picture, I don't think Anthropic can get away with this as much and I also think they're going to face a massive backlash from Max subscribers if they do end up ending the +50% promotion at the end of the month because Sol is a Fable peer and priced very competitively.
Then she hired a VA in the Philippines. Anthropic promptly banned her account without warning once the VA connected to the account. It took her weeks to get her account reinstated, at which point she had already moved on to OpenAI.
How can small companies with 1000x less money able to provide live support, but if you pay 20, 100, 200 dollars for a subscription you dont have a phone number to call ?
You message support, some real person reads and gets back to you within a day.
What your your company do? Is it low ticket business or a high ticket business?
Our free tier is time-limited but we still look at all tickets even from non-paying customers (in the hopes of converting them). A 1-hour intervention from a customer rep can result in a multi-year paying customer.
I can't tell you how many times I've experienced this with comcast. The last time I had to deal with it, was when I bought a new cable modem. I call in to provision it, the automated system assumes I have one of their modems and fails. For some reason I can't get technical support on the line and finally I resort to yelling 'cancel my account' over and over again until I finally get someone on the phone.
The guy was able to solve the issue in 5 minutes flat. The problem with automation is it's only ever going to be able to handle the 'happy path'
The solution here is removing corporate monopolies and political power.
They spend the money which can drastically cut into their profits.
you'd end up with a call center larger than most cities. It's not feasible.
Sorry, you can't buy an iPhone because Apple has got too many customers.
Sorry, you can't have a GMail account because Google has too many customers.
If firms had a base degree of customer support they were expected to provide, they would still exist. They would just not be as profitable, but customers would be better off.
I seem to remember there was a time when S/W was also designed with the aim to be easy to use, so that the need for support was reduced. It feels like the lesson learned was to keep costs low, not to ensure users were ok.
Right now we're using OpenAI's models by default since (unlike Anthropic) will allow us to use our pro subscription rather than token metering, but I've already had the joy of being able to change it to Kimi K3 (via OpenRouter) for an hour to try it out, and there were zero hiccups.
Using VPNs can also trip this.
I've seen some amazingly dodgy stuff when hiring people from south east Asia, sharing a paid account with friends worth a months rent there seems milquetoast in comparison.
Even if it was, it should be able to be sorted, maybe pay some overcharge or explain, and have your fucking business access re-instated.
Vibe coding is mostly garbage. But it can be useful for creating instant, disposable prototypes to investigate an idea or design direction.
Anecdotal example, I downloaded a new alerting app recently from an indie dev who had a small following back in the day in iOS space. It asked for a subscription, like $20/year.
I thought, let me try this the (final version, from Mac App Store) app first. Well, it's a barely-there vibecoded shit. There's a bare-bones list, everything looks like my nephew designed it, the macOS "app" is a iPhone-size view of the iOS one, it has a bug that if you click on it it opens multiple duplicates of the same list view for no reason that you have to manually close, and in general it barely works.
Yay for vibe coding.
The thing I miss most about programming is flow, and the constant bouncing between terminal tabs sucks. I’d love to do one thing at a time, with Fable, quickly.
An ide open with 20 tabs open each file a component, a class or an interface We also use to hold entire codebases in our brain.
Even a year ago when i was trying to do a hugo template manually with the help of the documentatin /tutorial, it was shit. The LLM at that time, was better helping me than the documentation.
I had some helpful people helping me on IRC / Quakenet.
But the hugo example i found very interesting because it was the latest hugo ducumentation and I don't think I was able to find a tutorial. I tried it without an LLM first.
Things just moved up on the abstraction-ladder, but it's still there, hidden beneath all the TUI sessions instead.
I'm struggling to scale myself even further. This tech is unreal and I have so many things I can do.
For the first time, tech feels like the 90's-00's again. Everything is greenfield and exciting and big tech is struggling to figure out what to do about it.
People are just hacking all kinds of stuff, and it's awesome. Feels like techno utopia.
It feels like the opposite of 90s - 00s: they were filled with periods where a person could self-study technology and get a job using those skills that few others had.
Where we are going (according to the AI-proponents) is children being able to replace you.
In brief; the 90s - 00s were a skill-valuation time, now we are looking at a skill devaluation time.
Unless you meant to say "Just like how any kid who could write broken HTML t put up a webpage could pretend to be a skilled professional, that's where we are now"...
Then they hit you with hourly and weekly limits, you are wondering when you are going to get cut off. Since LLM at probabilistic it often feels like pulling the lever of a slot machine, hoping our prompt is the jackpot. To make sure we win, we come up with systems, convoluted agents, context pipelines, rags to load . It feels like it's working, then bam you reach weekly limits.
(Just pay more if you want to keep winning).
I think developers need to wake up.
I myself started to use AI like a fancy debugger ,explainer. I make it walk me though every single line of code it writes.
I notice that i run into limits less, if i get cutoff, i can still make changes
The barrier and time between idea and usable implementation is almost zero now. I don't have to imagine. I can just write something and see it work before making larger decisions. I really like this. Many of my ideas were abandoned because I needed to study some obscure library. Now, I can learn the parts that I find interesting and just have the AI chew through the grunt parts easily. That's the good.
I started coding with a line editor on a small Casio handheld "computer" and used to keep programs in my head. I more or less knew what happened on each line without seeing the line. With larger programs, I had a mental model of what was going on where and a big part of the input to that was the effort of writing everything by hand. That's gone. It's not really important as far as the output of usable programs is concerned but there's a certain feeling of satisfaction that came with digesting a larger codebase and having it surrender it's secrets to you that's missing.
But for instance I used to work in 1-2 client projects at a time and they take months now I can do 4-5 at once and they take a month. That’s a huge improvement
Do you have project that needs to be maintained. I am also interested in your workflow. Do you Vibe code , never look at the code or do you hold the LLM agent's hand.
I feel like it's a spectrum
If anything maintenance is where it gets easier the mvp stage is where more focus is required
heh this is funny because this reaction was all the rage in the 90s early 00s too. You'd put together something you thought was cool and then post a link on a forum only to be told how it wasn't even "remotely impressive". I'm glad people didn't give up back then and i hope no one gives up now.
Command: claude agents
Which tells you everything you need to know about OpenAI.
Every time I get "you used your quota, come back in 3 hours, or 2 days" -> that is experimentation time with their competition, leading to changed service plans. When they said "claude -p" will be billed at API pricing even for plan users I moved my harness off claude. After I integrated codex, then it was never going to be a full claude project again.
What business encourages users to try their competition and adapt their usage to the competing products?
I've never run a local LLM model before. Certainly won't take as long to iterate on this.
I suspect the same will/is happening with AI. Either you will pay for it, or it will be so ad infested that it will become useless.
I have only 48gb of ram, so can fit only 80k context max, so good compaction is must.
I initially used the Claude Code harness on 3.6 A3B, but found that tooling would break as Claude released new versions and things would go weird. I've since written my own harness which has basic operations: read, find, bash (which can write files, python etc...) & web_fetch, all within a mac container. Works amazing. You don't need anything complicated to go very far.
Low hanging fruit would be Pi or OpenCode. If you really want a much better understanding of what your hardware is capable of then give writing your own a go.
Additional tip: Low Power mode reduces some token speed, but stops the laptop over heating and the fans going crazy.
https://support.claude.com/en/articles/15036540-use-the-clau...
Update June 15: We're pausing the changes to Claude Agent SDK usage described below. For now, nothing has changed: Claude Agent SDK, claude -p, and third-party app usage still draw from your subscription's usage limits. The previously announced monthly credit, which would have been available to eligible claimants in connection with these changes, isn't available. We’re working to update the plan to better support how users build with Claude subscriptions. When we have an update, we'll share it before anything takes effect.
If you are selling something, and losing $10 on each sale, you also would want to limit how much you sell.
I mean, sure, you are losing money on each sale so you can landgrab, but you still have to balance the land-grabbing with how much money you can actually lose.
I guess this is why they're pushing Claude code hard (not supporting agents.md, not allowing third party harnesses, etc) but when switching to another provider is as easy as opening a new terminal and typing omp/pi/codex your moat is effectively zero.
They can compete on price, quality or value but anything else is just madness. Currently they (arguably) own quality but this won't last.
They are high on their own supply. The people running these companies are delusional imbeciles who have been placed in charge of billions of dollars.
The breaking point for me was the privacy violation. They've been fingerprinting every request and violating users' privacy hoping no one would notice. Too bad, someone found out and that was the day when I cancelled my subscription.
https://thereallo.dev/blog/claude-code-prompt-steganography
Its in my opinion not wide spread at all and as today is literally the first time i ever hear anybody mention this.
The consumer side cheap monthly plans exist for the same reason companies like Cloudflare and Vercel have a free tier: When it’s cheap and easy to get developers familiar with the tools, they will push their companies to pay the real money for those tools.
It’s a hard balance with LLM serving because you can’t really make it free. $20/month is close to free, but the $200/month plans are in a difficult place where they’re big enough that many small companies pay for $200/month plans for their employees and ignore the enterprise features you get with the full expensive arrangements. So the companies are continually adjusting the $20-$200 plans to keep them from being reliable options for businesses, which is where the real money is.
There’s a short sighted cheering on of the 3rd tier and lower companies offering lower rates, but we’re already seeing them ratchet up the pricing and keep larger models closed after they get market attention.
Sorry, but people are cheering on Chinese companies (of whom your are unduly dismissive with your '3rd rate' comment given how good GLM-5.3, Kimi K3 are) not only because they are more economical, but also because they do not constantly refuse to do legitimate tasks and provide you with the weights for self hosting these models.
Maybe Anthropic's enterprise sales are going brilliantly, and the rest of us are just pixel dust to them.
Still. Brand perception is a thing, and between rug-pull usage policies, weirding verbedly output quality, and "I'm sorry Dave I can't do that" pushback, Anthropic are clearly having strategy issues.
After Fable 5 launched, it was better than Opus 4.8 for sure. Then they rug-pulled Fable from me (EU), and later released Opus 5. Now I only reach for fable when Opus 5 API returns 529 for the millionth time this year.
This is why I think open weight models will win out in the end. Right now there’s too much going on behind the scenes with the models. Day to day you never know if you’re going to get smart Claude or dumb Claude.
I'm strongly considering biting the bullet and just ditching my $200/month Claude Code plan for the Codex one instead, especially because I keep running into my weekly limits (even sticking to Opus.)
1. No 5 hour usage limit
2. Weekly usage gets reset CONSTANTLY. It's crazy. The longest I've ever seen it go without a reset is maybe 5 days?
3. I don't feel like OpenAI is constantly trying to fuck with me. Unlike Anthropic. I would way rather have Sol all day every data, consistently, than a slightly better Fable for like, 1 prompt every 5 hours, and only when Anthropic decides to not treat me like a cyber criminal. Believe in yourself as much as Claude believes your CRUD app is going to hack the pentagon.
4. Getting access to image generation, though I don't use it too much, is a nice perk compared to Anthropic.
edit: Should mention that I had like 4 banked manual resets as well. It feels like OpenAI wants me to use their product, whereas Anthropic wants my money while giving me a nerfed experience
This last reset took 6’ish days. I know because I was almost out of limit.
I have used Fable heavily on the lower Max plan and you are really exaggerating the limits here. I've had many multi-hour sessions with Fable on Max.
I’m sticking with gemini for now given the 20TB cloud storage and youtube premium that comes with it but I’m still open to switching to codex (which I’ve had a long term plus plan for). If google keeps delaying the pro models for much longer or makes them excessively expensive, I’m likely to switch out.
I used to have the Google One with Gemini and Drive space and YT Premium separately for my family. A credit card expiration lapse lead to closing YT Premium and being locked away. Why? it was because Google One plan bundles YT Premium Lite as an extra. So they actively blocked me from getting Premium back for a month.
Now I moved to YT Premium on my wife's account and downgraded One to lowest tier. Never heard of a company forbidding users to upgrade their plans before this. I suspect Google really wants users on Google One no matter what they want. I was barely using Gemini anyway, already have claude and codex plans.
https://codex-resets.com
The period was also marked with many billing bugs, like spending people’s usage credits for included Fable for a few hours (gave me a huge shock), but to their credit they refunded it.
Their disrespect for their users is also another problem. You only get one shot to make a good impression.
Same, it's been a while since I logged into Claude web.
ChatGPT web usage being separate from Codex usage limit is a nice touch unlike Claude.
With newer models text generation outputs have gone from probably human readable text to dense philosophical treatise about "load bearing" and incomplete sentences. So much so that now you need skills or another LLM to just parse the output. Simple answers and text generation just doesn't exist.
I'm not saying this with any undue derision, it's genuine - do they have a real product team or are they clauding that too? The direction makes little sense.
Instead often it feels like they make a change, then wait for someone to figure it out. Then Anthropic ends up being reactive as opposed to proactive in communication.
It is kind of funny because surprises from OpenAI tends to be positive (Tibo resets), on the Anthropc side I dread them.
The only B2C is going to be watered down ad-driven BS, and they will charge B2B via tokens.
The $50/mo - $200/mo consumer LLM subscription is not something I expect to last long / or to drive much of the revenue share... like individuals paying for Gmail vs Googles overall business.
I just put $30 on openrouter, switched to Pi, and I finally have a calm mind. Since I actually pay per request I want to maximize efficiency rather than utilization
From a cost-effective perspective the GPT models are much cheaper - one can easily tell it takes longer with GPT5.x to exhaust limits and this matters A LOT.
I can’t say which of these corpos I despise more though. I though for a while Dario was cool, but a massive distrust is piling and the first third player offering decent experience (and showing some decency) will win me over.
For the record - I’m also unsure whether I despise more Exxon or BP or burning fuel as a whole. Hope u get the point...
Also that the behavior of the models change, suddenly more fluff in the comments etc is annoying, but that is probably being part of using cutting edge tech.
(Eg. They repeatedly said they'd keep fable in lower subscription plans if they had the capacity)
Anthropic have been anything but. Flip flopping on model availability, model access behind an opaque filter, their past behaviour of model degradation as they prepared their next model… these are not signs of a reliable service.
I’ve mostly settled on using a mixture of open weights models through Together.ai and Fireworks.ai, a MiniMax subscription for high-token-use tasks that don’t need the best model (for $20 I get what feels like infinite tokens), and codex for the occasional high complexity task, although with Kimi K3 and hopefully soon GLM 5.3, it’s becoming increasingly less important. Deepseek 4 flash is my cheap main with delegation to other models as needed.
I’ve also found LFM2.5 8B surprisingly useful for single-focus tasks like “does this diff touch anything that isn’t related to the task”, and it’s incredibly cheap ($0.03/0.12 per M in/out).
But then, there are plenty of mindless, menial tasks out there, and it would go through those like a champ. And quickly, too.
To me, all of these are just exercises in getting me to pay for more tokens at API rates.
I’m at the point where I need stability and predictability. I want the B- student who shows up everyday rather than the A+ student that’s unreliable.
Most coders don't pay for tokens themselves. It's just on reddit and HN you would think that everybody does.
Work can pay a Claude sub. At home i see no need
In my company the average claude token usage is something like $5k/month/employee. Most hobby coders don't spend anywhere close to it.
Unlimited tokens aka unlimited cost is something that not every use case needs.
(And failing that, there was a real risk for OpenAI to be the default for enterprises.)
To defeat that, you need to frontline employees the chance to experience better tooling and models which is where the subsidized subscriptions come in.
But the agentic layer is being worked on, agents will start consuming more and more tokens
Their truth is “we don't have compute and are working to improve capacity”
People would root for that
Instead they got people rushing to escape the permanent underclass until they have a mental health crisis just to beat the fake deadline. $100, $200, is a lot for those people
That’s about A$16 a month in electricity if I ran it 7x24x30.
It doesn't help that all their models are bow trained to waste as many tokens as possible with their extremely verbose output
1) Set pricing tiers that do not change.
2) As models evolve, move the outdated models down the ladder, and replace the top tiers with the frontier models.
3) Give the users a warning before you do this, so they know their model is changing.
4) Sort the economic distortions out of your OPEX and reset pricing when the technology ossifies.
Is that so hard?
This keeps happening with every good new model release. People jumping from one to another, and as prices get lower, they start using the models even more.
People make not like to hear it but prices and usage limits are ways to shape traffic. The third option is the nuclear one like Kimi did, by just stopping to sell subscription at all. But that is something that DeepSeek can not do as all they offer is API.
Even OpenAI despite having the most compute is not immune to client influx = capacity issues. As people found their usage dropping, despite the push to the easier to run Luna models.
Reality is, that compute can not keep up with demand, especially when models get more capable and cheaper. What trigger people being using them more, what trigger compute crisis's.
This constant up and down cycle is going to keep happening for a long time, as this new market grows and eventually, somewhere in the future stabilizes. But yea, that is still going to be a few more years for sure.
I'm really sorry about that, but it jarred my ASD-ness
We all know that the subscription prices are not at all sustainable for these providers. You all do, right?
Yes, they're struggling to segment the market and find a way to make money, and that basically relies upon emptying the pockets of whales. As someone enjoying a hilariously subsidized Max plan, I understand that, and I don't think they're trying to scam me in some way.
And both sides of this equation understand that the market is competitive, and maybe more competitive than they thought it would be. Like, would you rather they did pull Fable when they first said they would? Or that they'd cut quota? I wouldn't. But I'm glad that Kimi K3 and GPT 5.6 Sol and the latest GLM and Qwen and...I love that this has forced Anthropic to change plans. I'm not going to hold that against them.
Yes, they are most certainly subsidizing the subscription plans. I mean, at least for people who utilize them to the quota.
But it explains why a nascent, hyper-competitive (I mean, clearly not remotely a monopoly) market has such erratic policies and pricing.
I don’t think it is fair to expect from people to know or understand this.
If you buy something or subscribe to a service there is a price tag on it.
You get X for Y amount of price.
That is how consumers conditioned for decades. They do not care what is your customer acquisition strategy. If Antrophic cannot provide reliable services on that price, customers will be unsatisfied.
The complaint wasn't "I paid for X and now they say they aren't going to give me X", it was "I paid for X, and they said hey guess what we're doing a promo and you get a bonus extra 2X, and also you get special limited time access to our new product Y", that's a hell of a thing to complain about.
Look, I pay a lot of money to Anthropic and I'm pretty happy that competitors have forced them to abandon their plans to add premium charges on these bonuses, but it's pretty ridiculous seeing the whining and gnashing, somehow turning this into complaints. It very much has a "oh no my lobster is too buttery, my blanket too warm" kind of feel to it.
I was with you until this part where the metaphor completely falls apart :p
https://www.pge.com/assets/pge/docs/account/rate-plans/resid...
P.S. batteries are the answer, hope this helps
Batteries aren't the answer.
I love how both of you are arguing about what the solution is, yet the problem isn't even clearly defined yet :P
Let P = batteries;
If (P == NP) then both are the answer.
But this seems apropos in a roundabout way since LLMs cost a lot of energy.
I don't think I can tolerate its writing style anymore. Reading Claude output is starting to cause actual psychological harm. I have tried many ways to get it to stop writing in its stupid punchy linked-in marketing-team voice, and I can't.
Is there a model out there that sounds sound this awful? It's like rubbing sand into the folds of my brain.
https://www.youtube.com/watch?v=71xxvp5R9hE
Fable is much better; still much too verbose, but at least I don't have the feeling that I am being charged for gratuitously added verbiage.
I have been changing my processes and the roles of my agents to avoid interacting with Claude Code as much as I can.
Sol gave me a straightforward bulleted list, easy to scan and read. Opus 5 though....it gave me a solid 7 paragraphs of how it worked. I read through it and yeah it nailed the same points Sol did but the output was way harder to read.
One of them writes better by default. One of them doesn't. That's the thing.
With enough steering, I can get a cheap low-capability model to do things correctly in most cases as well. But why bother?
You can tell it to write high quality code, to test things, and to come up with a proper rollout plan of a feature. Or it could just do it by default.
I know where I'm putting my money in that case.
I’ve tried so many things and I have to beat Opus 5/Fable 5 over their head every time. With Opus 5, the failure to actually improve the writing is borderline comical. Both models have been building up tomes of memories on top of my core rules, all to be ignored.
Nothing sticks mid-writing! The only lever is to ask to revise after the fact, a particularly futile proposition for anything non-trivial with Opus 5.
In contrast, GPT 5.6 Sol Max absolutely obeys my edicts to write well. I selected the concise writing style and its default writing is really not bad, but tighten it up with a standing AGENTS.md order and it obeys.
The downside is that even at Max, it’s not as good as Fable or Opus 5 xhigh at writing code. Larger work and the 256k context’s forced compactions cause it to lose track/fidelity of critical details. Review passes are essential.
Another reason I’m considering leaving Anthropic are the sporadic refusals… Fable printed a Markdown body with a hex dump to inspect for trailing white space and line endings, boom, denied due to `reasoning_extraction`. Had to tell Fable to never do hex dumps. Asked it a few times “what do you think is typically done for this?”, got another `reasoning_extraction` error. It forces you to lose a whole turn of work when you have to press Esc/Esc to “retry” the last prompt… except if that prompt was mid-turn, you’re losing your entire turn.
The recent BashFirst experiment is hella nuts, where it prefers writing bash over Edit/Write tools in auto mode. WTF Anthropic?
I pity those stuck with Claude in an enterprise/work setting.
As a side note, I think we are about too dismissive of computers and the Internet as “not real life”, that stuff written in text boxes by humans are just sticks and stones, etc. And now that we shouldn’t anthro-po-morphize or whatever the LLMs. But I do think, at least for some of us, that there is no way to avoid certain impacts that even non-personal text can have on us. It can feel alienating to interface with five different people in order to figure out how to solve a problem or navigate some burocracy (think Kafka). Well just interacting with one single LLM can induce that same feeling. The long-winded replies and the uncanny ways to miss the context.
Disclaimer that for my own needs I would be happy if the whole AI thing crashed and burned post-haste.
I really love CC and have been using it exclusively for a year now but reading Claude's verbose and semantically-obfuscated writing style is wearing me out and I'm planning to move to another provider.
I hope they fix this. It's terrible.
still don’t think anthropic models are worth the money
Often I'll tell it to summarize what it said only because it's providing too much detail, not that it's using esoteric language or weird claudisms.
No idea why people are still using Claude models.
My impression is they started on Claude Code and never tried Codex or other harness+model combos
Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline phone call or 1920's novel. So they _are_ going to train on them, no matter how many checkboxes you tick to stop them. Which is both a commercial risk and a massive new attack surface.
It's not just software developers; if law firms aren't controlling any public LLM prompts _very_ carefully, they can expect some meaty client confidentiality suits. In fact any organisation in a competitive environment should be worried: Acme Bolts: "Write me a presentation for Zoom Construction". Beta Bolts: "Is Acme Bolts pitching to Zoom Construction?"
Maybe this is part of why OpenAI and Anthropic are finding demand softer than they would like. And why open-weight models that organisations can run on exclusive hardware are thriving.
It's not even that expensive because even the 15€/month/employee bill for people who use AI very little can stack up.
Beside that though, many companies already have most if not all of their data in some cloud (Microsoft would be a prime example). Giving them a few extra bugs to get a (potentially) very useful tool is not out of the ordinary.
If you are under the impression that not going AI would threaten your business right now (which may very well be true for some companies) then there isn't really a choice even if you believe the vendors will steal everything eventually (which they totally will)
A large local MSP that has significant colo space just started a big marketing push towards hosted LLMs for their corporate customers.
I think its about to kick off everywhere.
It's more about the harness and API proxy you use. They need to be smart enough to know when a (self)hosted model is enough and when to forward to a SOTA model.
The SOTA model _can_ do everything, it's just expensive as fuck. But so is shoving a difficult task to a sub-par model that takes (relative) ages and comes back with the wrong result.
I have first hand knowledge of household name companies with billion euro IP who use OpenAI and Athropic products quite liberally. They have WAY more lawyers than I do and their sole job is to keep the IP safe. They wouldn't sign a deal with even a slightest whiff of the IP being used to train anything.
And if it happens, the penalty for breach of contract would have so many zeroes it'd be enough to buy a country.
It can be very hard to prove in court that your data was used for training
corp usually have all sensitive stuff ttl deleted for this reasons.
It very hard for a company to do anything without leaving some kind of paper trail behind that can be discovered in court. Not without crippling their own operations by simply refusing to digitise or write down anything.
Lots of companies filled with people working very hard to obscure their shady practices have been hoisted by their own internal docs. Just look at any major Uber, Google, Apple, Microsoft lawsuit. Do really think Anthropic and OpenAI are gonna be better at destroying their paper trail before the lawsuit starts?
It is a risk, but is it a big risk? If one of the big labs were to do this and get caught it would be suicidal due to the loss of confidence in them and the inevitable lawsuits that would follow for breach of contract. Given the AI labs are all desperately trying to paint themselves as Serious Businesses so that other Serious Businesses will pay loads of money for tokens the last thing they want is a rep for siphoning off sensitive customer data.
Because, almost certainly, it's not your source code.
This is the part that is truly scary.
Nadella the CEO of MSFT wrote that post “ A frontier without an ecosystem is not stable”
It seems to me that the unspoken assumption in that post is that no matter what happens they’re gonna be training on your data.
He’s the CEO of Microsoft, he knows how these decisions go down, he knows how the world works, he is sending a warning.
The frontier labs can have the models but without being where the workers are they cannot do much, lots of industries have strict requirements of not sending their data over to randos in the internet.
Microsoft and Google have the upper hand here with their workspace offerings and could easily position themselves as secure enclaves where you can use local LLMs where your data never leave your premises and is never used for training.
Did they? No, it worked out fine. Same with hosting your application at AWS instead of on owned hardware in a locked cage at the local co-lo data center.
Although of course the same argument was made against co-locating in a data center! Long ago an engineer carefully explained to me that no serious business would put their data in a co-located data center since the data center operator could just plug in a hard drive and take it all.
It turns out that contracts do actually mean things, and businesses want to do business with each long term. If you’re at a tier where Anthropic contractually commits to not train on your data, I would just sign and move on.
Still does not address ecological concerns, of course.
MS ain't some altruistic company having core mission the good of humanity, as they proven across decades.
Yep, all they've got to train on are bot-generated Instagram posts. Can't say I pity them, though.
Complete nothingburger. Actually worse than that, its a London Horse Manure crisis.
>Destroying millions of obscure books to scan them?
I really don't see the issue. They got slapped in the face for trying to do things the right way and torrent the lot. Why wouldn't they exercise their legal right to buy physical items and create digital backups?
>They make meth-heads look scrupulous.
My local meth head checks in on my family every 2-3 months, because when she had fled from hospital post surgery, and added some meth to some morphine, we gave her new clothes and a safe place while we convinced her that the ambulance service wasn't run by Satan. She's good people. Anyway if she wanted a whole bunch of digital books I would help her torrent them like a responsible person.
I find this argument weird because you or I are below the threshold of being targeted over book torrents these days.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
I just kept a $20 plan going for use on my phone.
> "The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives."
From: https://www.anthropic.com/news/claude-fable-5-mythos-5
> "Some attacks only become visible across multiple requests. Best-of-N jailbreaking, for example, sends hundreds of slight variations of a prompt in the hope that one will work. Larger patterns of misuse, such as state-sponsored espionage or data extortion campaigns, only surface when our safeguards classifiers can zoom out across many requests. Detecting these threats requires temporarily retaining prompts and outputs so they can be analyzed together, rather than one at a time."
From: https://support.claude.com/en/articles/15425996-data-retenti...
I don’t doubt people are hitting it… shrugs
Nowadays Codex handles the bulk of the implementation and Fable/Opus on the planning.
Not sure if Anthropic patched it, but early on its release the web UI Fable guardrails will trip if you mention you're a biologist.
Used to straight out bail to Opus on this, now it's "Honing" and "Pondering" for 10 minutes on it. Then I get a bunch of "This response didn’t load." and it still churns ahead.
I have the same experience. I found that if I add to the prompt: “but it’s not security related, we only focus on the auth library design”, it “bypasses” the safeguard.
One of Anthropic’s problems though is that the legalese will say one thing, while key employees say something totally different on HN or X. At the end of the day the lawyers always win.
2nd time was literally me being lazy and telling it to commit and open a PR.
Very strange.
(yes, it was a very lazy prompt I could easily have googled, but that makes the refusal even more bewildering)
I do not know why people even use Max or Ultra levels of thinking effort. Most of the time, i found its better to just run any of the frontier level models, with high or medium. Their capability beyond those levels is often diminishing returns, in exchange for double or quadrupling the costs (or usage).
And the bonus is that often at those levels, they tend to spin up less sub-agents that just eat away at tokens/usage, like its water in the desert.
A Max or Ultra is something that really only belongs if medium or high can not solve a very nasty bug or issue. Even in planning its often too much.
I did a quickie test with Fable when they were OMG TRY IT NOW, didn't see much of a difference from Opus 4.8. Except i ran into one of their stupid security guardrails. Yes I'm doing security on the software my employer owns, silly.
Then I didn't notice the US unbanned Fable after banning it, or that they extended the "trial" Fable period where it worked on fixed price plans. Not in time to bother trying again.
After that Opus 5 showed up. Language changed, okay, I don't care much. However it seems to create busywork for itself. Spawning subagents on tasks that don't do that with 4.8. Overall taking longer.
So I'm back on pinned 4.8 and doing actual work. Main problem - for Anthropic - is it's good enough.
Where can I find these stats?
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
Reverting is annoying but being flagged for security questions while I'm doing code review is insane.
You want to disable "Switch models when a message is flagged"[1]
[1] https://support.claude.com/en/articles/16049681-why-claude-s...
So stingy. Even with OpenAI's recent usage troubles, they're still so much better than Anthropic it's not even funny. Re-ran the code review with Sol as a benchmark and it turns out Sol's performance is within 70%-90% of Fable's. Anthropic's still got the best model, but what does it matter if I can barely use it?
I'm on the $200 plan (work pays) and I also have the $20 OpenAI plan (I pay) and keep a balance on OpenRouter.
There is nothing as good as Fable, not even close.
I recently had it run a 18 hour autonomous rebuild of a project (moving from Spark to Pandas for performance/data size trade off issues).
It orchestrated Opus sub-agents flawlessly for 18 hours. It even did a great job of managing the number of agents to keep them within the 5 hour budgets (I think I had to restart it twice).
After 18 hours I ran a /simplify, /code-review, /simplify cycle which went for another 6 hours.
2 billion tokens (mix of Opus and Fable), 24 hours of continuous coding and a bug free outcome. It would have cost $2000 at API prices and worth every cent.
Fable's ability to keep other models on track while working on these long horizon goals is so much better than anything else.
Far from neutered, I've never had a cyber refusal, and Fable's English is actually readable (unlike Opus 5).
As an aside: while I hate reading Opus 5 English it still is a noticeably better model than Sol in my experience.
But I could handle losing Opus5 is I got Sol instead. But there is nothing even close to Fable.
Codex is more token efficient and tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents.
I spend 10+ hours a day in both agents, typically side by side. I often have them do direct “bakeoffs” from identical prompts in separate work trees. Most of the time, Sol’s work is better than Fable’s. Not always. It’s situational. But it’s certainly not the case that Fable is in a league of its own or anything.
This is very true, especially vs Opus 5.
> tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents.
The strength here was less the code quality and rather the long horizon task tracking.
This was a very large task - I was chatting with the maintainers and we estimated 4 to 6 months work over multiple phases for human coding.
Fable is able to handle that long goal, with incremental steps along the way, handle the verification and course correct when it finds a problem.
I think the larger context helps here some, but the strength of the model on this specific thing is notably better.
I'm not alone in noticing this. https://www.primeintellect.ai/research/nanogpt-speedrun shows Fable is able to manage a run nearly 1/3 longer than Sol (8.7 days vs 6.1 days). In my use cases Sol is much closer to Opus 5 though.
Sol in my eyes is powerful, but it over engineers so much, that its actually a liability. Where as Opus 5 is slightly under develops but you then can give it a small push for what is missing.
I rather have it under develop and i as the human in the loop, can correct/enhance it. Vs the models that adds so much, to the point that your going "dude, stop!". Remember, removing code for a LLM is way, WAY more difficult then adding it.
My main issue is with Sol is that its designed to over engineer without thinking why its doing something. Great that you security harden 1000s of lines of code, but ... nobody will ever get to that code. Its that lack of intelligence is where the model becomes a issue for me.
Its easier to have less code and then do security audits, with you approving what needs to be changed/hardend.
It may simply depend on the developers their mindset. Some folks just want the models to do everything for them, and performance or code bloat means nothing to them (forgetting that this bloat over time makes future LLM work more expensive).
Its funny how everybody has their own opinion for what model is better, when in reality its more about that model fits your own development style better.
That's 1.3M tokens per minute. I suppose you mean input tokens? This would not have cost $2000 at API prices because you would have cache hits.
I think the rewrites are the main story for llms in code (hot take). Writing greenfield code at the seams also something which might work well.
It's a fairly large code base split across 3 repos.
The good thing was that it is fairly easy to verify: we have a working (but slow) version that uses Spark, with lots of existing unit tests.
We verified by using those unit tests as well as running our end-to-end process in the Spark and Pandas version and verifying the two databases were within the differential-privacy noise bands of each other.
in your instance, anthropic may decide, arbitrarily, to stop 'autonomous rebuilds / refactors and ports' because they could pose some alignment/rights/etc risk to whatever slop their philosophers dream up while they're out eating $200 avocado toasts. then you can't do the thing anymore.
fable is good, absolutely. agree it roasts Sol which is, comparatively, a little receipt-hunting jack**
but now imagine being an enterprise, and having another organization not only taking your workflows and baking it into your models, but then deciding they can arbitrarily cut you off.
when you can instead own your data, use an agnostic provider, and get better results (through model combinations), it will take 1-2 quarters to figure it out.
the main reason anthropic is killing it is because they really do understand the enterprise development experience and lifecycle and have built products and have a sales-team that can deliver.
business-model and vibes-wise they have lost all goodwill in the past 6 months, and that momentum will be quite hard to regain.
feature not bug
a bit of lost revenue is an okay price to pay when you are familiar with getting bullied in hs
I was chatting to the maintainers on 2 of the 3 repos this affected and we estimated 4-6 person months work of we were hand coding.
The PRs on those 2 repos are 48,000 lines of code (which is a problem in itself!)
I'd estimate that generating the specs would have been maybe 2 weeks work? It's across 3 repos, and I'm only really familiar with one.
> done in 1/5th of the time with smaller models
The coding itself might have been faster, but the end to end time would have been much longer.
> Maybe try using your brain.
Believe me, my brain was a load bearing seam in this task,
This is true, but only for certain tasks. Even as a Fable fanboi, Sol is much better at Fable for some non-programming tasks: Fable for life-planning tasks is miserable because it keeps adjudicating rules, where I've found Sol to be insightful and warm (characteristics I'd previously associated with Anthropic models).
We're probably less than a year away from all the frontier models being so good at everything for day-to-day use that it doesn't really matter which you use, which is going to seriously fuck up the business models of all of these companies except the infra companies.
I should have pointed out that I always prefer Sol for OpenSCAD for example.
so violating the TOS? Or work pays for a plan you cannot use at work?
I was doing work at work on a work task using a plan paid for by work.
Anthropic know that lots of people are doing all sorts of “bad” things like employers paying for Individual plans, (and using multiple accounts to get more usage) and aren’t yet enforcing the rules… but by the letter of the Anthropic terms, your employer should be paying Anthropic a whole lot more (and that’s one of the reasons why AI usage is going to get very very expensive as soon as the subsidies stop, you and a lot of other people are already paying a lot less than you should)
I think multiple plans are against the ToS, but I'm not doing that.
The Teams plans are more convenient for a number of reasons, but yes, they top out at the 6x plan, not the 20x plan.
Edit: ToS are here https://www.anthropic.com/legal/consumer-terms and https://www.anthropic.com/legal/commercial-terms
I've re-read it and I'm pretty sure there is nothing that forbids a business paying for a 20x account. Notably they say this in the consumer ToS:
> If you use an email address owned by your employer or another organization, your Account may be linked to the organization's Anthropic enterprise account, and the organization’s administrator may be able to monitor and control the Account, including having access to Materials (defined below). We will provide notice to you before linking your Account to an organization's enterprise account.
which goes at least moderately close to indicating using it in a work environment is allowed.
Not this again. There is not a single shred of evidence for this. In fact multiple times this year alone, people from Anthropic have said that inference and deployed models have positive margins. The big bucks are always being spent on training the next model.
I don't think that means the subscriptions aren't subsidized though! I expect they are, at a carefully calibrated rate.
They want to maximize people trying them.
The real money is the enterprise plans which need a seat price plus API rates (unlike the Teams plans which are a pretty good deal).
It's a simple fact that paying enterprise token rates would cost many times what the individual plans cost. It may well be that the enterprise income outweighs the cheaper tokens on the individual plans for net profit, but using the individual accounts as subsidy is a common tech industry tactic and what you said doesn't prove they're not doing it, either.
It's a little hard to believe they wouldn't be trying to slow the burn rate for an IPO if the inference were truly so profitable. And i find anything they say a little hard to believe all the time anyway (though admittedly they're a lot more trustworthy than OAI)
HN's take that they "kill their own company" seems wildly out of touch.
Truth is, Sam and Dario and Elon are all terrible and so are their orgs. Leaving any one of those companies for any of the others is wild.
Zero Data Retention, for the uninitiated readers in this thread.
The no-ZDR is clearly to permit surveillance. I would be shocked if NSA wasn't all up in these SOTA model providers' systems.
Never subscribe!
It's to train better models. The three letter agencies don't need to spell it out in a ToS, they just access it if they want.
This is in the article.
The AI companies have already shown how little they care about other peoples intellectual property and they won't care about their customers either if it stands between them and the promises they made to their investors.
ZDRs are backed by contracts negotiated against extremely capable and well-resourced counterparties. Being found to violate them would bankrupt even OpenAI and Anthropic many, many times over.
To the extent that you view these companies as your adversaries, it’s worth improving your mental model of their actual incentives and constraints.
What I don't see is vast areas of industry finding $10s to $100s of billions of value in LLMs. There's no lint or compiler that can check for correctly constructed contracts. So LLMs, which should be useful to law firms, incur a lot more manual checking of their work than coding agents.
Less formal document production in other industries is likely to have less structure. That might not matter in some settings but I'm having trouble thinking of an example off the top of my head.
Using them for coding makes it easy to self check its work (assuming those pieces of work are "verifiable").
As proofreading will still need to happen, what do you think the appetite for lawyers is to do this kind of work? Do you think it will drive fees down significantly? Empower younger lawyers at firms who probably are the ones doing this checking for the partners? (Or will that just create a further divide).
I'm genuinely asking as I am not in law but all my family is and it's nice to see someone here that's thought about the impacts in that space.
Will it drive down fees significantly? Doubt it, there's not enough pressure on them, and the industry is resistant to change. Firms don't want to make a big deal about using it because clients will then ask why they're not getting a discount.
Will it empower younger lawyers? Not as much as I'd like. I'm very fortunate to have an employer that lets me use Claude Code for my work (to a limited extent). I think for 99.9% of lawyers it's not an option available to them, through a mix of concerns around AI usage and concerned IT departments. There could be great benefits, but it would rely on having to break out of traditional private practice which would make it difficult to get enough work.
I think in the near term LLMs will have much the same impact on the legal industry as it has on the software industry.
Funny you should say that. I read somewhere (not on HN, but I think it was a post linked from here) that a number of law firms who deal with extremely sensitive documents have started buying amped-up Macbooks with 512GB of memory to be able to run local models.
These are businesses who literally - and for once this word fits - cannot afford to let some of those documents get anywhere outside their corporate walls.
There was a thread earlier this year about the option getting pulled from selection (https://news.ycombinator.com/item?id=47296302) but I'd guess you can still top the thing up yourself.
Coding agents are great because they have compilers, linters, test cases etc to ground themselves in.
With tools like OKF I’m sure most knowledge work would be distilled to its core data - it’s AST if you will, and then allow models to guard against hallucinations.
Checking if a case law exists is a tool call, you can demand provenance, it’s all _buildable_.
Hallucinations are “solvable” this way, so the rest is just time and adoption…
I compared insurance quotes last week. Needed cover for 2 brands, 1 company. Opus jumps up and down saying both brands need listing on the policy schedule. Human broker said not.
I told opus and it's the usual "thanks you're right" bollocks because it bothered to read in more detail and found that all business activities are covered.
How do you think the industry will respond to contract clause slop ??
I can foresee each side inserting 100s of innocent clauses with minute dependencies that provide hidden advanatges
There are definitely linters and this exists https://catala-lang.org/
But that doesn’t mean that an appropriate harness (even Claude code and a dynamic workflow) can’t complete, and validate, an answer.
That's not really the point though; it doesn't have to be "legal specific".
The key is in understanding that "Basic chatbots (claude, chatgpt, etc) hallucinate. Independently verify the output (even if it's another model doing the verification) before assuming it's correct".
Ask any LLM about a random game mechanic and then be prepared to verify if the answer is for the patch in 2025, 2023 or 2020...
The only way we got a head start of using it for coding and maths was to have some formal method of validating the output as part of the training and inference time.
Translation between languages.
That value dwarfs all programming value that can be had. Economically, culturally, scientifically, spiritually.
Now do a full movie's subtitles with google translate, in an automated fashion. So that it understands the context and still translates it correctly.
Just by nature of how often it's used etc.
You need workloads for AI be cost effective: software, automation etc.
Hackers always rage and down vote every time I mention this, because they are unable to see beyond their small world. Why didn't they learn that their part of the internet is 0,000000001% of what the world uses the internet for today. It's going to be the same with LLMs. Programming and hacker stuff is going to be 0,0000000000000000000000000000001% of what the world uses AI for. But translation is going to be in the top 5 of use cases.
A developer using sub-agents will consume more tokens 1 Day than a marketing manager will consume 1 Month, easily.
Unless there is something inherently automated about the nature of the AI, it will be a tiny % use case.
Even a lawyer, using AI daily for contracts - that will be relatively light use. They'll make more use doing legal research etc.
Developers and Automation are the 'primary' uses cases for AI, and in the future, we'll start to see AI integrated into Apps - that will be 85% of tokens consumed.
Yes - once translation becomes realtime, and we have our Star Trek Universal Translators, then translation will become more visible, but even by then, a relatively small part of overall consumption, even if it's more highly visible.
And a programmer uses one million tokes, which gives him $100 in sales (or value).
What is then the value of a token?
The comment I answered asked where there is an industry finding 10s or 100s of billion of dollars in value from LLMs. The answer is translation. It's the value they as customers get out of the LLMs, not what cost they are paying for the LLMs.
Value has to be counted in production, not consumption.
> Developers and Automation are the 'primary' uses cases for AI
Just like programmers and scientists were the primary users of the Internet when it began. But things change rapidly.
The main problem is I think you're assuming because the current translation market is large (I'm just going to assume it's ~100B in size just from a cursory search), then it will remain large with LLMs. If LLMs are much cheaper than humans, even with a lot of growth in translation volume the total spend may not compensate for it (again most of the volume will probably be using almost free models?). Another is assuming that because something is valuable you can charge a lot for it. Like, oxygen from air is extremely valuable to us. If oxygen somehow depleted we would die almost instantly. It does not mean everyone goes around purchasing oxygen or even less that you can charge absurd amounts for it.
Before AI translation became available, I would hire professional translators. Their rate was about $50-100 for a detailed product page into one language, and it would take them a few days to deliver.
Now with AI translation, I pay about $120 per year for unlimited translation. Meaning dozens of product pages into 5 languages, plus e-mail back and forth with hundreds of customers. All instantly at my convenience.
What this means is that a whole lot of people, sectors and businesses who would never hire professional translator can now have high quality translation at their disposal for a cheap price.
> Another is assuming that because something is valuable you can charge a lot for it.
Even if LLM translation won't deliver trillions of dollars in income to the AI companies, it will without a doubt deliver trillions of dollars in value to customers and users within the coming few years.
Good human translators are still higher in quality than any AI and will always be. But AI translation is currently far beyond good enough for all use cases, except fine literature and maybe complicated juridical stuff. But I'm not familiar with those sectors.
But we'll see. Text editing tools are included on all digital devices, yet companies pay for commercial solutions like MS Office. Cameras and basic image editing tools are included on all smart phones, yet there are millions of advertising agencies around the world. And so on.
eg. processors are needed for everything on the planet. No Intel isn't going to generate annual revenue in the trillions because of that. In the case of frontier companies the moat is even smaller with dozens of players competing. LLMs will become a low-cost commodity. You meanwhile can make billions building applications using them though.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
I doubt it. The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription. Even if my #NeverSubscribe movement succeeds, all that misdirected hardware [into datacenters] won't likely be practical for home use.
The only time it really succeeds is when regulations force it. I can store files on my own computer or a home NAS just fine, but if I start a medical practice, I have pretty much no hope of being HIPAA compliant without signing up for a data hosting service. The same goes for tax preparation, banking, and several other fields.
It may very well be the case in China, Europe, or Australia that regulatory restrictions force users into models-as-a-service, but in the US, regulations censoring models, even if they frame that censorship as a safety measure, won't pass constitutional muster.
I had not considered that as a possibility. It is somewhat dark and unlikely, imho, but a real possibility.
I am more inclined to believe that the fab capacity will grow over time and "commodity" compute will be available to us all again.
However, I can also imagine going back to the 60s era "hyper verticalized" mainframes, in which case, the frontier labs might not just be producing models, but chips and an ecosystem around themselves and their suppliers/customers.
If you're actually willing to pick up old DC inference hardware on the cheap you can already do a lot at home. The new hardware not so much admittedly but if the Chinese models keep getting better and running on less hardware... I don't see a reason to be quite so pessimistic.
- 10-1000 person groups such as corporations where you can amortize serious hardware with parallel use
- Porn.
Sometimes I'm just trying to sus out if I'm truly seeing things these days or going a little nuts :)
I don't know why...
Is it A/B testing?
Is it load shedding?
Is it because I'm in Canada?
Is it because I'm not on the Claude Max plan?
Is it because I'm not paying via API?
Is it because I'm not paying via Bedrock?
Is it because the U.S. is worried people are distilling?
Is it because the U.S. wants to keep the top capability to themselves?
I think open models are the future. Anthropic is killing their reputation so fast. If they don't come clean I think they're cooked.
So what you are paying for may vary on a day by day basis, which is quite undesirable, even if their main goal is simply to make a better model. When it comes to a tool, I'd rather have consistent mediocrity than instability.
Me: why did you add HTMX?
Opus 5: you asked it twice.
Me: quote the exact sentence(s) where I asked it.
Opus 5: I can’t because you didn’t.
Agree that Sol is a great model. I find that it's improved a bit but I mostly attribute this to its eagerness to use the harness' memory features. (I'm using it in Hermes Agent, FWIW)
I don't think it's something deliberately malicious like planned obsolescence but it's more like startup culture of "just make it fit in this sprint".
Your explanation is probably the most likely and largest contributor. Anthropic states that Opsu 5 is a "pinned snapshot". They claim the weights and model configuration are not silently updated, BUT the surrounding serving infrastructure can change, including the request router, safety classifiers, and sampling logic. Anthropic has stated that if behaviour unexpectedly changes on a stable model ID, an infrastructure update is the most likely cause.
Further, "High" isn't a fixed amount of compute. Anthropic describes effort as a "behavioural signal" and not a token budget, with the model deciding how much thinking to do. So their "High" might be "Low" now, and we would never know.
Finally, I strongly suspect some quantisation or KV-cache compression is happening. Anthropic doesn't clearly delineate whether this would fall under the pinned weights and configuration, or the infrastructure, which almost certainly guarantees it's the latter. Forgetting earlier information, poor retrieval of details, contradicting previous conclusions, hallucination, degraded instruction-following, and losing the thread during complicated tasks are all symptoms of quantisation and compression.
And they're somehow shocked users aren't loyal to that. It isn't about cost.
The bar is low.
I don't want Shakespeare, I want Bob the builder.
Half of my work is telling claude how to behave. I'm pretty certain they have enough _data_ to realize people do the same thing time and time again.
Check this comment of mine for a better explanation of this: https://news.ycombinator.com/item?id=49413353
But people are not as alike as you think. I doubt I share your unique preferences.
That said, I don't spend much time telling it how to behave. Are you sure you're not fighting the default system prompt?
Dijkstra in the Foolishness of Natural Language Programming
[...] the "naturalness" with which we use our native tongues boils down to the ease with which we can use them for making statements the nonsense of which is not obvious. It may be illuminating to try to imagine what would have happened if, right from the start our native tongue would have been the only vehicle for the input into and the output from our information processing equipment. My considered guess is that history would, in a sense, have repeated itself, and that computer science would consist mainly of the indeed black art how to bootstrap from there to a sufficiently well-defined formal system.
Imagine if only we had languages at our fingertips whose explicit purpose was to precisely and unambiguously tell a machine what to do!
https://www.cs.utexas.edu/~EWD/transcriptions/EWD06xx/EWD667...
This work was fraught with bugs, a large portion of which came down to disagreement of what was coded v/s what was written. Even if you had airlines sign off sentence by sentence exactly what you wrote down in English, that's too open to interpretation.
People don't appreciate how the same sentence can be read five different ways by different people (or the same person on different days). We had to structure our documents to be closer to pseudo-code than to English to get any meaningful consensus on the definition.
But that would require people to think! And the marketing is they don't need to do that...
Last consumer (ish!) product that required people to learn something to use it was PalmOS with Grafitti, wasn't it?
I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.
I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic
Also, Sol doesn't refuse constantly and speaks like an engineer rather than a deranged academic.
Opus 5 is legitimately terrible and can't or won't follow instructions. It is of negative utility and does more harm than good to my codebase.
And - trick - give it a little cli so it can run Codex if you have them both.
Let it do codex to do the bulk of the work, get a OP5 sub-agent to audit the work of the codex worker.
Just let Op5 manage and have 'specific oversight.
You can run for 2 days on 1 context window in the manager, the advantage is that it will stick to a broad plan.
I have Sol xhigh drive Claude via tmux and I get amazing results until I run out of Fable. Then Opus comes in and starts acting like some sort of autistic academic with OCD.
It tries as hard as Fable, but isn't smart enough to do it well. It starts designing ever more elaborate tests, frameworks, and procedures while making up rules for itself and piling them on top of each other until nothing gets done. It's the ultimate bureaucrat.
Worse - More than once it's spent days in a loop because it invented constraints for itself that it couldn't satisfy then lied and told Sol that the user imposed those limitations. I don't know if it's actively avoiding real work, or just isn't capable enough to work the guardrails that were obviously forced into it.
Even fable is way below human level at many tasks. If fable were as good as you claim, all computer jobs other than "frontier scientific research" would have been replaced by now.
Most of the benchmarks have exceeded their usefulness. Opus 5 beats fable 5 on many of them. Anyone who has used both models will notice immediately that this doesn't translate to the real world. Opus 5 is nothing short of a regression from Opus 4.8. Fable is genuinely a great model so long as you don't trigger a guard rail and it downgrades.
Sol in my experience isn't significantly different than fable ignoring that Sol burns usage 10x faster but the end result is hard to differentiate.
GLM 5.3 is a hair behind these two.
An anecdote but not an original one from the people I talk to.
But Fable security false positives and pricing just make it not worth compared to Sol IMO.
Either way, a model used to solve the top 10% of problems that people use AI to solve for, being used 10% of the time … seems like it’s in a decent place.
I find Fable indispensable, and measurably better than alternative models, for complex feature development in an existing codebase. It’s the closest thing I’ve seen to nearly one-shotting features. Even still, I only use it for the hardest features and Opus 5 does a good enough job on the rest.
It is not only that every single turn with opus on high reasoning (high is the middle setting) takes at least 5-7min. It is barely usable interactively. Instead of a chat it feels like you're sending emails to it. Tasks that used to take an hour when it "reasoned" for 45s before it started doing anything now take almost entire day.
At least until few weeks ago it was horribly, mind bogglingly slow (a little better during US nighttime), but the quality was still good. I could not do things interactively, but providing prompts were fine it built stuff fine.
This is no longer the case. It makes stupid errors all the time. So you cannot leave it to complete some work, for example write infrastructure migration scripts a night before then you simply run the scripts and perform the migration during the day. Nope, every single script has stupid issues requiring use of the model to fix them. As they are written in it's own "spaghetti code" fixing them by hand is not an option.
It is clear to me they are doing some shenanigans behind the scenes to try and optimise their compute use. Either they quantized these models dynamically or do other things that affect quality.
In top of that they now do this stupid fingerprinting.
Anyone who knows how output vectors are turned into tokens knows it will eat up a lot of compute or destroy quality.
What if:
- Opus 4.6 was the last Opus generation that got a lot of use by Anthropic's own employees
- After that they primarily used Mythos internally
- 4.7, 4.8 and 5 were RLAIFd by Mythos "teachers"
- Hence why 4.6 is the last Opus gen who doesn't report back like a robot wanting to cover every potential hole another AI system would've spotted and criticized
- Hence why coding style in Opus 5 also gets criticized, not only behavior in CC
Worth noting that Fable (ie, Mythos) is actually nice to interact with.
[1] https://github.com/zachahn/vomit
I switched from Opus 4.7/4.8 to test Kimi K3 a few weeks back and the test hasn't finished; it's my daily driver now.
Given their general behavior and preference toward social engineering to scare the shit out of normal people...this seems fitting.
This is just a case of competition working: Fable (and Opus 5.0) just aren't as good as Anthropic believes, and there's no switching cost, so everyone's moving away.
Once the government stepped in, their 5th generation was effectively killed. They either had to lock it up (Mythos), neuter it (Fable), or leave people with the perception they were overpaying for a weaker model (Opus 5).
Hopefully Anthropic has learned to anticipate this risk and has a plan for rollout of their next model that plans for capricious ad-hoc regulation.
Option C: they could "get hacked and the models exfiltrated" plausibly. Then the horses will be out and there would be no more point for the USGov's barn door.
Among other things, "we have stolen the collected works of your culture, now help us grow so that you can all lose your livelihoods and become our serfs" has to be one of the absolute worst marketing approaches in history.
Crimemaxxing about to go off the charts!
(Or perhaps vulns that NSA/etc discovered and has been keeping it private; as they're known to do).
I feel there's a lot that's unaccounted for, and the whole "AWS team reports a 'jailbreak' that is just 'review this codebase'" story doesn't add up.
I wonder if there were some parallel construction going on, and if at the same time, the NSA started losing the exploits they had because it was getting patched.
While there is some fallout likely with much more effective hacking being easily accessible, I’m sure govt analysts (unless they were let go) have their own prediction models telling them it’s inevitable that this technology eventually makes it to everyone they don’t want having it, what with China seemingly releasing every progress they make openly. Which makes me think that they’re preparing for that inevitability by hardening the govt systems currently in place and/or by burying the secrets they want to keep hidden deeper underground.
Given the history of the US and this particular administration, I feel burying things deeper is a greater priority.
-t. American
Model prep costs far too much money to operate under this kind of regulatory regime.
A government staffed by logical, sane people would have figured out how not to encourage this, like mobilizing the IT and security industries to secure their products faster, while also cracking down on the "our AI is the most dangerous tool in the world" rhetoric bandied about by the frontier companies. But we don't have that, we have a government of Loony Tunes characters. When you put extremists in office you get extremist behavior.
But this isn't what happened? It was only a couple of months ago! How can we be getting this completely backwards already!
Fable was aready what they were selling to the general public, and that is what the US government stopped them selling (not Mythos!)
Anthropic tuned the already existing classifier to handle the use case the government highlighted and then they went back to selling it.
They actually loosened the restrictions on ML-programming using it too.
I use Fable as much as I can and I've never had a cyber refusal for it.
Mythos was locked down to a few customers.
Fable went wide.
USG told them to stop selling Fable.
They tweaked Fable to appease the USG, but it clearly refuses/downshifts more requests than e.g. Opus 5. (I have run into Fable refusals even doing really normal stuff like fixing bugs reported by static security scanners like Brakeman. Not downshifts, outright refusals.)
So the net result is they put "Fable" back on the market, but it was perceived to be a worse overall experience, at much higher cost. (Remember all this happened before most users had really even had a chance to put Fable through its paces.)
That's where we are today: Mythos is locked, Fable is neutered. They probably can't fix this until they are ready to release something that's not called Fable, or that they can say has some key architectural differences to Fable. The sword of capricious regulation is always going to be hanging over Fable.
> You're agreeing with me.
No I'm not. They are selling it.
I used Fable very extensively before the ban (100% usage in every window available) and continue to use it now.
I haven't noticed any change in the refusals before vs after.
> it was perceived to be a worse overall experience,
Sure, people will believe whatever they want to believe. Doesn't make it right though.
Computing technologies are relentlessly deflationary. If the value of their TAM wasn't a fraction of a manual process they replace, they wouldn't have a productivity advantage. And the amount of TAM per unit often declines over the life of that product category.
I would be unsurprised to find investment in data centers to be 10X what was really needed. The same goes for where the LLM S curve starts to flatten.
this is about selling credit to clueless boomers
But a while ago, I had access to a Fable harness that just never gives up. And that verifies itself. It burned $100 in API tokens in 15 minutes ... but it succeeded for all the prompts where Fable + Claude had failed.
And I believe that's a real issue for Anthropic. Fable+Claude is not too expensive thanks to the subscription, but Claude severely nerfs Fable. To save money, I guess. Fable API + Custom Harness is a different class, it's so much better. But API tokens are so expensive, you're cheaper off hiring a freelancer.
I would compare it to a extremely high end $15k PC, or an expensive pro-grade video camera, or a freight train, or a …
I would say at least 95% of the global population will not encounter a situation once in their life where it would be actually useful/warranted.
It’s more reliable and makes less dumb errors than Opus.
It still messes up, of course. But for my working style, I definitely prefer it.
That said, Opus 5 is broken. Use 4.8 or another vendor for the build agent.
Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.
Cheaper, more powerful AI will continue to expand the bubble. Projects will get more ambitious. Everyone will build out their own custom little software. Code diversity expands and requires even more AI.
They are already good enough at what they mechanically are.
You have to use the right harness, right verifiers (automatic where possible, human where not), etc much much more specific than a generic one like claude code or codex, and it will also be able to work within constraints and be the "proposer" of an imaginary optimisation problem and an excellent one at that. But you have to frame the task at hand in that manner or maybe even reorganise the task you do itself so it is more amenable to being framed that way. If you use it this way, it is _already_ massively economically useful. But it will take many years for it to actually be usable in that way, since you need DC capacity to come up first which is few years away and also well, massive organisations that have to integrate these will usually take many years to do so.
It is also useful albeit less so in cases like general SWE, where you still need a human in a loop for non-verifiable requirements, and also in other general usecases where information retrieval is too intractable and you need to carefully use LLMs as a component of the overall system.
I am not saying Fable or whatever the biggest models are are useless - they will certainly be useful for tasks at the frontier of the day - which is today complex exploits and open math problems, and well, tomorrow it could be something in biotech. But this is not what the entire bet is on at all - just automating day to day drudge at the tens of thousands of massive companies and governments we all know and love is more than enough. With the right training data (which _also_ is a bottleneck and takes time) you could even automate certain processes entirely. Sure, if we get a crazy medical innovation and end up saving trillions in healthcare great, but that's just a bonus.
None of this is to say that I think there is zero sketchy financial engineering going on
In recent months we’ve had the first automated unmanned amphibious assault, the daily drone count in our hot wars is jumping by leaps and bounds, and arms suppliers are promising future drone shipments in the hundred thousand unit range. The ten year picture for reactive combined swarm intelligence on the battlefield is promising to be widespread, highly lucrative, and in need of constant adaptation to near-peer efforts. Datacenters in space are dumb, datacenters in space to power orbital weapons networks and rapid response capabilities make sense.
On top of that we have international trade, scalable customer service, and a first pass 80/20 answer for businesses focused elsewhere. Shitty, maybe, overpriced, maybe, but useful enough our grandkids are gonna use ‘em.
In both cases, as well as potential new LLM-like tech, there’s an argument to be made for being a leader now to dominate the future. That means compute and tech positioning, and memory & GPU deals.
YouTube was a money loser, Google was ‘losing’ money on them for years, YouTube didn’t have a sustainable business model. YouTube was the biggest, though, and whatever premium Google paid to be #1 then meant they were #1 when the online video business model matured. Now they’re printing money with a platform outcompeting news, social, and video platforms.
One line I remember was that he said the disruption is not happening the way they thought after GPT4 due to inertia bla bla and that it will be slower gradual change and they got the timelines wrong.
And yes on the defense usecase. That is the main reason governments are giving a hoot about AI. Orbital datacenters too, I know people working on the Indian one, it's entirely for defense usecases. Basically for missile stuff.
And you can have extremely simple work that you nevertheless have to do as a professional. For example, drafting a routine customer email, summarizing a meeting, formatting a report, filling in standard documentation, or making a trivial code change.
So I don't think the $15k workstation / professional camera analogy really holds. Those are specialized tools whose capabilities are mostly useful within a fairly narrow domain. A general-purpose AI can be useful across thousands of completely unrelated tasks, including one-off problems encountered by ordinary people.
I have tried fable, gone back to sonnet, tried to use agent mode to mix and match and various other combinations. If I decide that Fable is the only one that can save me hours or days of time, I will go back to it and pay the cost. If the cheaper or free ones do well enough, I will stick with them.
When you see the price of hardware needed for these models, we are generally not paying very much towards that price so I expect prices to start ramping up as the honeymoon ends and these companies' investors get itchy for their ROI.
Fable is not used because it has extra retention requirement that corps can't sign off so it stays disabled for everybody in many cases.
It's also not winning on day to day work against Opus 5, which is simply available as there are no extra retention requirements and no extra paper work to do with legal.
Devs also don't like that Fable refuses to work on half of their prompts.
Then comes better pricing and on-par capabilities from competition.
In what universe? Hell, Opus 5 seems to be worse than Opus 4.6 at almost everything.
This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...
When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use Fable still *feels* much better. I think benchmaxxing the new open weights models is ruining the value of benchmarks, if they ever had any. When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. And considering the point I'm poorly trying to make, that Opus doesn't feel like frontier anymore, this is probably the last month of my Claude subscription.
I've wondered if this is part of why we don't see the reasoning traces for Anthropic's models before -- Open models might just be accurately surfacing how the sausage is made.
Anthropic in particular is much more compute-constrained than OpenAI and SpaceXAI and has relied on partnerships to provide inference. This reality factors into their pricing and usage limits (they started 'adjusting' the 5-hour limits during peak hours, and it certainly wasn't an upward adjustment). Accordingly, this is presumably what Anthropic wants, given they develop and release the lower-end models, suggest users use them in various nudges within their product, position the bigger/more expensive models as "For the most complex tasks" in their UIs, and so on.
I think the problem (for Ant/OAI) is that there is no sensible lockin or moat. LLMs are essentially interchangeable and stuff like a harness doesn't offer enough value on its own for someone to be locked into using one of them.
Now with the onslaught of the Chinese models that offer almost the same quality for much less money they have a very serious problem on how to proceed. Investors now might be looking through rose tinted glasses but their patience has its limits.
Enterprise is a whole different ballgame IMO with countless technical, compliance, and employee adoption considerations that add friction to switching. It's also where both Anthropic and (as of last week) OpenAI get the bulk of their revenue, and, incidentally, the venue where US Government regulations on Chinese models would have the most impact.
At this point, every model writes code about as good as I'll need for the stuff that I'm working on, and with the right harness, I'm perfectly happy to let them cook, until it comes up with something working. Whether it takes 10 minutes or an hour really makes no difference to me.
The frontier models may be better, but who cares if the last generation of models are plenty good enough for what you want to actually do?
And the best local models are quite competitive with those older lab models.
If you are disproving Jacobian Conjecture it makes sense to be on SOTA, but for writing Golang and Typescript, faster sol/fable/opus class models are imo more likely to get user interest than the latest frontier.
1. They acted as if they were so far ahead capability wise that they could stop listening to their users.
That is basically it. It is an extremely common belief that Opus 5 acts like a condescending wannabe-thought-leader, yet Anthropic's response to this is largely been "You're using it wrong. Try deleting all your config files".
They released a "concise" output format, but pretty much swept it under the rug, despite it being one of the loudest complaints about their models. It also, generally, does not work as described.
Not only did Opus 5 get much more difficult to work with, but it got, from a customer view, significantly slower. The token rate might be the same, but if every interaction takes 50% more tokens, it's 50% slower.
All this time OpenAI has released a slew of new models, lowered the price on them (which, bluntly, 80% of user count don't actually care about since they're on subscriptions), increased subscription capacity, and increased response speed.
It's not about cost. It's about Anthropic being the frontier-lab version of the marathon runner who decides to celebrate to early, and then loses the race.
Nobody has complained and seems like for every use case we have Opus is more than powerful enough, especially with Opus 5
I find it funny how OpenAI got caught lacking for a very brief window, but it turned out to be a very critical turning point.
Like a guy that that's at the top of their game the entire year, and the one day they have the flu, the CEO does a surprise performance review.
Sure there are Microslop and Oracle db users but most of the world we live in is Postgres and Linux. That's why I think most companies will run llm's like that.
> good at its job just use it instead of re-inventing the wheel
Exactly why is everyone reinventing a harness every month. There will be Microslop / Oracle harnesses and there will be 1 or 2 open source ones that win.
You can't make it an indispensable dev tool if devs cannot rely on it. They will absolutely jump the service to s more reliable one and that was Codex for me.
P.S. Claude is indeed down right now for me.
Also Sonnet 5 was released June 30th, seems to be grouped in with 'other'?
https://ramp.com/data/ai-index
(click on model market share)
Saying "people are not buying that many Rolls Royce, instead they buy a lot of normal cars" is kind of duh. Anthropic is clearly still operating at inference capacity.
If I need to write all the specification so it follows it, I might just write the code or use a cheaper model.
Opus 5 and Fable 5 have been quite disappointing.
I used sol, terra and Luna and I think they and the first two are good for first code reviews. Luna is not much better than deepseek V4 flash
This wouldn't be so bad if it didn't make mistakes periodically, especially when it's about to tap out. I get the sense that if I upgraded to the $200 subscription it would get me a lot more usage, but it would still run into these issues anytime I sat down to work for a few hours.
I'm just using medium effort, so it's not like I'm on high all the time.
Then on top of that opus slurps tokens like Anthropic is afraid they’re losing money. Read 200 lines of file, +20k tokens. Excuse me?
if people dont see why they need a model this smart, they probably arent using ai enough
I am not doing awfully complex tasks though. I imagine a lot of other people are in a similar boat, either switching from Claude to ChatGPT or even just min-maxing DeepSeek V4 Flash 0731 or similar.
At this point, Fable is really just an experimental model (as it should be). It can do very useful things, but is it a broadly good general purpose model? Definitely no. Most users don't have problems hard enough for Fable outside of coding very large and complicated projects. Most users don't have 45 minutes to accomplish a task Sonnet can do well enough in five. There's not a PowerPoint in the world where Fable is the right tool to build it.
The secret sauce is going to be in letting Sonnet decide to delegate to Opus and Fable when they're the right tools for the job. But you can't train a model to do that until the bigger/better models exist and you can observe how your users actually take advantage of them. I'd bet money that's exactly what the rlhf going on at Anthropic looks like right now.
Economically, it makes sense. Sure, on paper you want users burning as many tokens as you can. But pushing users into burning tokens and taking a long time and getting a meh result is far worse then giving them the "fast and good enough" solution that occasionally burns more tokens automatically when the problem requires it, and getting a higher quality result out when you do. From an infrastructure capacity perspective, this is the dream: you stop measuring cost [for Anthropic] per token and measure cost per outcome, allowing you to use less hardware to accomplish the same abstract units of work.
DeepSeek v4 Pro did same for $1.7 in 7 responses (peak-off times).
For doing much complex task, I would be super nervous about using Anthropics models.
For AI agents, we have been mostly using GPT 5.6 Terra.
Did they not learn? Performance is a feature
and resources or tech stack tips from HN?
I've been running Qwen27B Q6 with 130k context at 30tps. It's not bad.
I think with an "autopilot" mode in pi, and sub agents to handle context window management better, I'd envision you could get close to unattended workflows. I haven't quite gotten that far though.
Negative is you have to do it all, and the temptation to tinker is real.
I can respect the guardrails - I also can see why OpenAI may not have much control over their models - but I need an AI who will do whatever I ask and not play judge and jury.
I mean, who says "screw you" to requests to get 35+ year old vintage computers working? Claude, that is who. Its guard rails are so stupid. I hear people trying to do simple mailing list management hit it too.
I am just about done with them.
It's unclear if Mythos2 or 3 or whatever they're calling their next model will be an improvement for most common enterprise use cases.
LLMs can't solve basic things (writing non-slop documents, understanding context without massive handholding) and for coding other models are quickly becoming 'good enough' without the same cost and nannying.
That's why Anthropic is 'stealing' workflows.
But it turns out it's much harder to push adoption when your users don't really want to use your product.
Code was a unique use case where the code luddites were loud but a minority - most people don't want to update 300 cases of variables across their code base for a name change. Most don't want to write unit tests.
There are a few use cases where that will happen (law is next, maybe quant finance) - but otherwise most companies are throwing money into a pit and getting 0 return.
It's a very interesting race and state of affairs, but Kimi K3 and likely the next DeepSeek models will put the high price token affair to rest.
Unless of course, mythos / next model really does solve some universally applicable problem that people want it it to do.
I've now bought the $200/month plan from OpenAI and simply ran `/status` in Claude to get my session ID, then asked Codex in the same directory to "Please take over the work started by Claude Code with session ID `[…]`." This seems to work like a charm. It also seems that the Codex weekly and monthly budgets are more generous?
Unless Anthropic adjusts their customer-abusing behavior I think I will end my subscription with them soon.
I become more disillusioned every day.
It seems like, by having the models write Python code, they tend to write Python code like an average developer. Which is to say, quite bad.
Add in the complete failure of the models to adhere to instructions in Claude.md, memory files, and added multiple times in prompts, I’m wasting huge amounts of time fixing bad design decisions that the model just slips in.
- Fable, nerfed or whatever, too expensive and not fully included in subscriptions.
- Opus, neurotic (excessive) slop machine.
- Sonnet, way too token hungry, cost much more than 4 series.
Their almost daily outages does improve the experience. Anthropic really have messed up this year.
Bro if you want us to spend more on fable let us use our whole damn rate limit for it. Enterprise is a different ballgame obviously but for the subsidized Claude code users they literally cap it
https://www.youtube.com/watch?v=wTiYaWFP59Q
The whole "we're going to replace all of your workers with an agent while also making AGI happen" thing was a bad idea in the first place; one that could not have gotten anywhere outside of Silicon Valley. People just want tools that they can deploy at scale, not to be your beta testers for the singularity.
Most of the people pushing this are just hoping that they can cash out before the hype pops and financial gravity crashes the party. Sam Altman recently claiming that the singularity is here is so stupid on its face he should just be treated as what he is, a huckster.
None of this stuff ever made any sense on what it was being sold initially. It was always insulting that the media and business leaders tried to argue that the tech could replace entire call centers or vast swaths of entire industries.
People keep arguing, but it will or it has based on extrapolating certain, reasonable use cases. Klarna has shut up about replacing call centers with bots because Markov chains with memory only can do so much.
I’m happy with Claude. If they become (bigger) jerks, I’ll switch to something else. I don’t ha e the energy to praise Anthropic today and I won’t have the energy to demonize them tomorrow. The emotional investment people have for/against these companies feels like celebrating or being offended by the weather.
TBH I don't think any of that is unique to the IT industry's relationship with the Valley. Other technology-driven industries have a similar worship-ish relationship with a few rarified businesses. But the culture of the IPO exit accelerates all the most short-term motivations to do anything.
Anthropic is the fastest-growing software company in history. They have no issues "attracting users"
No shit we're all paying 清冲 Flash to do the grunt work. Turns out though, paying 清冲 Flash a few more cents does exactly what Ivy Wasp Pro Mythical does. Crazy how that works.
I can literally open a new chat with just "Hello" and it gets bumped.
Claude keeps track of memory if you've turned that on so all conversations are somehow tracked over time Claude randomly will mention that I'm a developer while I'm asking unrelated questions and say oh because you're a developer you might like this or because of my background and infrastructure you might find this interesting and I always get creeped out by it.
They have the concept of incognito chats, but I can see those being worse without context from other conversations happening.
In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview.
Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro.
It does have issues with availability - but it's on a free promo right now. That also means that I don't know how much it would have cost if I had to pay API prices for it, which might not be cheaper than the subsidized small company / consumer usage, but for large companies paying API prices for either offering, the difference will be considerable.
https://openrouter.ai/stealth/ox-alpha
So yeah, I and I guess others, are quite active in whatever little way we have available, to up vote new models and share stories.