Hacker Newsnew | past | comments | ask | show | jobs | submit | armanckeser's commentslogin

Sounds like free market at work to me. I am just glad there is quite a lot of competition in a field that I would have assumed would have huge costs of entry

Building small models is a lower barrier. Think of something like to just train on a corpus of internal corporate data. I've see some small models that do this. It's like a super RAG thing. I think more of that will happen. Excited to see a SLM vendor emerge.

I wrote about how to redirect to nitter instances on any platform a month ago [1], and a few days later we had the cease and desist. Fortunately, you can redirect to twiiit.com and that will find a working instance for you.

[1] https://armanckeser.com/writing/redirecting-x-com


Why assume token costs will increase as opposed to become commoditized?


Self-hosting takes either cloud- or on-prem-engineering chops, which are in short supply. In my view, that becomes the primary driver for lowering token costs. Secondary are the alternative model sources like Z.ai and Moonshot, which between open weights and lower prices help to drive commoditization. But there is a political risk with these, for example, Z.al is on the U.S. Department of Commerce Entity List, which blocks American companies from selling advanced technology and goods to the firm.


you seem to be discounting or unaware of the possibility of a) automating a significant chunk of new buildouts, and b) domestic companies seeking a slice of the “sell open model hosting for less than competitors” pie.

I truly believe the cost per token for most people will eventually go approximately to zero even in the absence of subsidy unless there is a dramatic state intervention to curtail free computing and hardware R&D and manufacturing.


Fair callout. Strongly discounting (b) and (a) is underspecified for me to respond reasonably.

Your premise of full token commoditization has a major wrinkle though. GPU-centric data centers with never allow tokens to go to zero since you have a 1.5 year technical deprecation and 5 year life on the GPUs. Maybe Cerebras or similar chip builders, after writing model weights to silicon, will improve those returns on capital and enabling the scaling you are envisioning.

From my perch, AI at the edge is severely underserved, and the GTM for the hyperscalers mostly ignore it. So if someone figures out how to scale AI at the edge (nvidia + HF perhaps) then I think your premise is on point.


Maybe they still want to serve the feee tokens if that gets them good training data.


On macs you just install a desktop app. What can be simpler? Decent Qwen models already doing some simpler types of agentic work.


Sure but if the customer installs the app on their own computer how do companies charge them a recurring subscription fee? Adobe can get away with it but most cannot.


That would be even worse for them, since then current customers of these AI-middleman corpos would just start deploying their own self-hosted and private LLM services. In before the complexity argument, there may well be some completely different middlemen corpos, who will assist in such deployments and initial ramp-ups.


Right now they're losing billions/year, it's not sustainable


Because that's been the playbook of every tech company for a decade or two. It should be our expectation that this stuff will get an expensive rug pull the instant the companies can get away with it.


> the instant the companies can get away with it

LLMs have no lock in. Switching costs are zero. It's really not clear that instant will ever come


Yes, this exactly.

People in this thread are discussing it as if it is a question of pure economics. It's not! It's a question of "how much do we think we can get away with charging"


There seems to be a weird self-fulfilling narrative here:

1) "AI companies will go bust because tokens will cost too much" >> 2) "Don't let them build data centers that will sit empty" >> 3) Demand outstrips supply by a lot >> 4) Token costs too much.

This is no different from the NIMBY's that block all the housing then sit in a huff saying "see! everyone's leaving. good thing we didn't build anything!" when the rents go up.


Talk about a goomba fallacy. The people who are against the data center buildups (which is increasingly becoming a bipartisan issue the world over thanks to the antics of the AI psychos) are not the same people complaining about token prices. I'm sure many of them would even be gleeful if the AI companies went belly up


Doesn't matter if they're different people or if it's politically popular.

It's how the cycle would go, and in the end the same people will complain about how AI "is only for rich people" (how high token prices will express itself in populist language)

The exact thing happened to housing.


Again, the people who are anti-AI don't give a shit about token prices. They will care that AI is only for rich people, but this won't be due to token prices, but because only the wealthy will be the ones with any kind of job (if the dream of the AI overlords is to come to fruition anyways) or income stream due to displacement caused by AI


"dream of the AI overlords"

I don't care what their dream is, I care about what's happening IRL. Please read the original article again, AI is clearly being used by workers to expand their productivity (and their earning capacity). The main thing that will screw that up is if token prices spike.

"AI is only for rich people, but this won't be due to token prices"

Statement doesn't make sense. Higher prices means unaffordability. I don't know why that would change for AI tokens. Higher token prices is what will make sure advanced models are only available to capital owners

This feels "unintuitive" only because it's not the dominant narrative, but that's not a measure of accuracy. "Everybody believe so" is not a good defense.


I am not sure the author of the comment you are replying to understands that LLM systems have prompt caches


There are many reasons why one might not want an account, and I presume your view is exactly why X created this gate in the first place. However if you are like me and your main use case was say Android, a simple app install is basically no work. I wish I were able to do it on the network layer without a CA though, that would have been the silver bullet


If your aim is to just get started, get yourself a Claude account and just download Claude Code. You don't really have to read anything on it, as you use agents you will get to learn what they are good and bad at but defaults are quite good anyway if you use a capable model like Opus 5. Truth is there are a million guides and plugins in the open but they really provide incremental gains for a beginner as they would overwhelm you.

If you are more conscious of the price, try out opencode or pi and research OpenRouter or similar for coding with models like GLM 5.2 or Kimi K3 or even Deepseek 4, they will be less out of the box but cheaper.


I have been using OpenRouter for a while now, and my Uni is self hosting Kimi K2.7 and Qwen 3.5, so I do know my way around chat completion. I've tried Claude Code with Kimi but it just felt confusing to use (probably cause I didn't know what I was doing) and the results were bad (probably because the codebase I've inherited is a mess, and it can't comprehend the web server logs actively lying (and neither can I)).

I've heard Claude Code alternatives can be better for non-claude models, but again, which ones? How can I tell if a tool is good or bad when everyone has the same LLM'd readme? How can I avoid security nightmares like OpenCode when they seem so popular?


Places like r/selfhosted try this and to be completely frank, puts that community into the "naive" side of history. One day not so far in the future AI assisted/generated will mean more reliable, better tested, better documented by default, and do you think AI companies are not focusing on making their LLMs not sound like AI? In 2 generations everything we know about identifying LLM generated content will be obsolete. Only thing that matters is if the content is quality or not.


Tangent but what a horrible website at least on my android. Starts as light mode, suddenly shows three streched columns, turns to dark mode and the title is hidden behind the header. After years of being around how can they end up with something like this


Whenever I'm bored, I go to the ant-design site to see how long it takes to find the first gross accessibility violation. This time it was the dropdown just above the fold, which wasn't keyboard controllable.


I bounced when I saw the three broken columns but never even saw the dark mode switch. It’s that slow.


Completely broken on iOS as well. Feels like AI slop.


It's existed for a lot longer than LLM generated code has


It's gotten much worse than I used to remember it, though. It used to be pretty bad before too, but it at least had an ok homepage.


It was always far too buggy for my liking as-is, though it’s amusing that somehow that’s become worse over time it seems


It is funny that people now instinctively think of AI when encountering slop, as if the slop that the AI produces was not deeply informed by the slop it trained on. Plus, the quality bar on the web has always been low, masking the true capability of LLMs.


I personally find AI doomsday preaching incredibly arrogant. You think that we will create super intelligence, something that will exponentially get smarter, and the thing it will do is... end humans? Why wouldn't it just find solutions to problems we struggle with? Also people don't live to work, its asinine to argue AI replacing jobs is the problem, obviously the problem is the gains being concentrated higher up. Another thing is, we argue AI needs to be maximally useful to everyone in the world, and the argument to do that is government control??? I would rather push for AI to be superintelligent faster than any one human can control it. Frankly human controlled superintelligence is likely a worse outcome then uncontrolled superintelligence, at the very least there is no guarantee one will be better than the other.


I don't use poetry anymore but do check the updates before claiming such things

https://python-poetry.org/blog/announcing-poetry-2.4.0/


Fair, I didn’t do a “as of this morning check”. Should’ve done better. It’s sad because I moved away exactly because this feature was missing and now I’m not going back.


May 3rd 2026. Release too new, didn't read.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: