Hacker Newsnew | past | comments | ask | show | jobs | submit | AM1010101's commentslogin

Awesome, keen to see where this goes. Can we have light mode for the whiteboard and dark mode for the code?

definitely down to expand the theme set! could you file an issue for tracking ? Thanks!

Did 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?

It did not. xhigh is new to grok.

Not sure on the API side, in Cursor you can always use 4.6 at xhigh.

We've only used it through API - but you're right, now API supports xhigh for 4.5-4.7.

I think 4.6 got an xhigh after launch. The benchmarks seem to all have been against 4.6 high.

Would the idea be to have many specialist models (forms, wikipedia, final cut, etc) and have a parent model choose which is best (potentially also a specialist model for selecting specialist models) with the idea an llm would give a larger goal and trigger this cascade of specialists to quickly do the task?

I think this is the logical next step in AI. There will be a few large "smart" AI models, but mostly smaller more hyper-specialized. In large part because they are faster to build....and faster/cheaper to run. As the model this post is about shows - you don't need to be able to write the works of shakespear to be able to run a computer.

The frontier models are already getting STUPID expensive. More than the average person certainly can reasonably afford for any use case. And even the business users are having a hard time justifying the prices for anything other than the most bleeding edge, "big picture" things (like creating a huge project plan).


Conscious thought seems to follow this too. No one thinks to breath, we often get in a groove and do deep thinking during a rote task, etc.

I mean this is what MoE is the first step of. Slice your intelligence to make it more efficient without losing much capability.

CUA seems to disaggregate. Train specialist models and make delegation explicit. But I think the future for capable but efficient models involves doing it all in one package. I have no idea how that architecture would look like or how to train it.

Like a human would, I can walk and talk, without thinking hard about walking and still not stumbling. Specialized sub-circuits that interact. Not via a text protocol, but by directly interacting.

Did you see the Astra plays Portal 2 videos? You can see how separate viewing a scene, understanding the scene and acting in the scene still are. And in this model of interaction, a very capable model still makes stupid mistakes.

It is remarkable, that it works at all. But it is neither effective nor efficient. And in my opinion the real revolution comes not by capability, but by integration.

So to tie the loop: I think CUA makes an interesting first step towards specialization, but I am not sure if that is the logical conclusion of how AI should work.


This is a pretty good move for Meta to go all in on. AI is ChatGPT and Gemini to normies, Meta is not considered. However when I ask one of my less technical friends if they connected any tools (email, calendar, docs, etc) they are clueless that this is even possible. To them AI answers questions and generates images. If Meta can get this to the masses as the first AI thats useful at doing agentic things, ie an assistant rather than a chatbot, then the average user will see meta as as the company who provides this to them. It’s their best chance at winning the category and having a strong play in AI. If I were Chat or Gemini I’d get something to market fast. Timing wise are they also hoping to beat / be compared to any siri announcements this week in the media too?


I think the average user doesn't have enough going on to need that, or they will be unable to articulate what it is they want the agent to do.


I don't understand why google are being like this. This is an opportunity to extend an olive branch to developers (like openai has) and a great way to get data around how people are actually making this model work for them regarding coding…


Seems to do reasonably well in opencode according to artificial analysis. https://artificialanalysis.ai/agents/coding-agents

If I had to pay per token I would probably consider using this (they seem to be on the pareto of performance) but not being able to use opencode with a subscription is not really something I'm realistically going to do when claude and codex are around. Also never gotten along well with gemini-cli / antigravity-cli.


Interesting how quickly they seem to have caught up.


50% off at open router is also still applied so it comes out at $2 / $10 per 1M.

Feature request for Artificial Analysis, allow us to see these live prices on the pareto. It would amazing to also see what a 25,50,75,100 % utilised subscription costs compared to raw tokens.


Not true. I have used Fable for all my planning work in a reasonably large but well organsied codebase and I rarely hit a session limit. I would use all of the Fable specific usage if I accidently let it do the full implementation of the plan. Pro was perfect for me. I would then used GPT-5 / Composer 2.5 for implementation followed by a Fable check over. I will definetly see if the API usage in Cursor is enough to cover my Fable needs before upgrading.


Plannig and review. You want the best plan and when its done you want to catch all the bugs. Implementation can be almost anything.

Of course you can do the planning and reviewing by hand but sota models do a pretty great job, especially in review (please still read the code).


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: