It is kinda funny I am always interested in Anthropic's new releases, but i have a great distaste for their actual release posts/videos. They LOVE these stupid style videos for any feature they release, their text posts are usually just a ton of nonsense that i dont wanna read, and now there is this huge "revelation" that their models can pump this annoying format out.
I say Bummer if more people adopt this style, its gonna become less and less authentic feeling as the months go on now.
Its the same across the board for everything, no one has an answer for the noise AI creates. Really cool tech, but the timing is so bad because we just can't even fathom what dealing with the surplus of generated garbage looks like.
The answer is likely moving away from the internet entirely and having actual face to face screening right away. no company wants to commit to that, so its gonna continue to be some weird expectation from both sides where the applicant doesnt wanna be screened with AI, and the hiring team wants X amount of effort put into every initial application.
I think it just must move entirely offline, until some semblance of an answer exists, because right now like i said, we can't even fathom what a solution looks like online.
if its remote initially, closed eye interviews where you just have a discussion might be easiest way. Pretty easy to tell if someone is bullshitting if they can't discuss topics that likely should excite them somewhat, or at the very least topics they've been around for years.
Feel like any test nowadays should INCLUDE ai, but you likely just want insight into how they are getting from A to B. I don't wanna assume what you mean by cheating with AI, could be as simple as answering live questions with cheating software or whatever, which is obviously a huge issue. I wanna respond and say it should be expected for any take home type stuff, and be rolled into the expectations if anything, but that could be way off of what you mean.
both sides of the entire process is just a nightmare now, one side has all this noise to sift through, and the other side is pressured to mass apply because of all the competing noise they can't be seen through. I was listening to something about AI generated responses, and they seemed to have no empathy for the people who gave human responses, but used AI to give them an overview of the company and what it does for the question "why do you want to work at x?", and sorta mocked the person because they can tell they asked for a generic description of what the company does.
There has to be some give and take, initial applications if you are spending hours researching a single company, you are setting yourself up to be absurdly let down if that same company sends you an automated rejection within a day. A reasonable person will only do that so many times, before just giving up and waiting for the other side to be the first mover instead.
I've been doing this, my only experience with codex was brutal usage wise and i just retreated back to pi pretty quickly so the credit usage i'm receiving is kinda all im familiar with. Surely seems like less than CC, but i guess not using codex makes my experience kinda not valid for comparing usage.
And ya i can go over that 240k limit, I still very seldom do, and try to treat it as the actual limit. I'm surprised to see so many people still talking about compaction to complete long running tasks, i think the bulk of the work should be somewhat frontloaded into a plan that is split off into subplans, then you can kinda open up a few options, one session with subagents for the subplans of the main plan, or just handoff prompts about progress against the main plan/relevant subplan. I just never trust the blackbox that is compaction, I feel its a recipe for disaster/context poison.
Ya all these articles lately about how everyone is sick of reading AI prose, and interacting with models in general. Tons of new model optimizations and workflow optimizations or whatever. I'm not really aware of any idea or product aimed at making the internet usable, and making it somewhat resistant to the generated noise. I think HN is a bit better than reddit for this type of example for floods of comments, first movers on reddit REALLY rise to the top and stay there.
This is a preexisting model being optimized. Its absolutely not some unexpected release after that blog post. I won't defend that blog post, but saying THIS release is proof they don't mean they are slowing down is just incorrect, this is a prime example of what i consider horizontal improvements
Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.
The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.
What does model naming have to do with pacing or not?
This is a ~20% relative quality improvement on the frontier (fable) at ~40% of the cost, just 21 days after the last release.
Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.
pretty annoying topic tbh. You're just weaponizing this dumb blog post so anything released is now a contradiction. By your same logic, if all inference was served at 50% less power cost and the savings are passed on somewhat to the user, its also a contradiction of the blog post.
Its an agenda serving blog post, but constantly bringing it up like this is just obnoxious.
Well yes, If you magically found a way to reduce a models energy usage by 50%, the thing that would happen immediately after is a doubling of a training scale and of the test time compute assigned for a given budget.
I don’t see how that wouldn’t be considered pushing the frontier. Given that scaling is the one thing that has been bringing us closer and closer to AGI, a sudden 2x increase in scaling laws would definitely not be considered pacing. You seem to see a contradiction where I don’t.
what's absolutely obnoxious is not being able to have a normal discussion, where we don't need to resolve to useless strawmen that are a loss of time for everyone.
No, releasing a model that has ~3-6x the performance/cost ratio than your last release just 3 weeks ago is not the same as making one employee 0.001% more productive, but you knew that already.
No one said they can't push the frontier, it's pacing the frontier, which mean very different things.
the 5.0 name means its the same model, they made it more efficient and less horrible to talk to. The top frontier people are all still talking to fable 5.1 or models not released to the general public, yet you want to claim this is pushing the frontier, instead of evening the playing field. I state power efficiency gains for existing models and you also claim thats pushing the frontier . It as hell takes a lot to get you to say they AREN'T pushing the frontier, which is why i resorted to the extreme strawman of desk location, because it doesn't seem they are allowed to do anything otherwise.
Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can "do it", but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
Cool maybe this makes the 20$ sub less of a joke. I deemed 5.0 unworthy of spending time fighting with and personally considered it by far the worst release of 2026 by either of the two major labs(promising model, but obviously not even close to ready for the general public).
So for that 20$ tier for the entire summer and into fall, i was on their 2nd class public model(4.8) released in May. Not surprisingly it became my grunt model, doing the simple work. By far the least I've used Anthropic models in the last 2 years.
yep, all i want is freedom to make a workflow where i don't feel tied to one provider, CC is exactly what i don't wanna get trapped in. I'll happily use claude MODELS, but if it means i have to keep learning two harnesses side by side to keep using one particular provider, then the second i can easily replace it, i am going to(even if its a small drop in performance).
I've really been liking the car analogies for AI lately.
Personally i dislike driving, i dont like cars or trucks or traffic, yet i own a car. I can protest cars and they wont go away. We could shutdown the 2 biggest automobile makers and it likely just messes with prices/availability for the general public, but doesnt put the genie back in the bottle, hell we can shut down all car manufacturing and it is the same lack of shutting down driving, it just creates a lack of surplus supply.
All these things are pretty true about AI as well, nothing puts the genie back in the bottle.
What we can do, is try to strive towards electric cars, stop putting lead in fuel, enforce laws so driving becomes safer etc. Same with AI, strive towards greener electricity, find some sort of solution to deal with the absurd noise problem that is pervasive across every single area any generated content touches. Build some structure so intelligence is accessible to the average person rather than only the privileged/governments.
One of the worst things about AI, i think is the timing of it. We have built our modern internet experience into auto suggestions, and have no way to filter out what is and isn't a real person's content, or even higher effort content, and all those suggestions have perverse incentives to GET suggested in the form of ads/clout/whatever. Pair that with the absurd noise problem generated content produces everywhere.
I say Bummer if more people adopt this style, its gonna become less and less authentic feeling as the months go on now.
reply