Hacker Newsnew | past | comments | ask | show | jobs | submit | more returnInfinity's commentslogin

Brad Gerstner confirmed that tokens aren't being sold at a loss. Whatever the formula, API + Subscription split, the companies are making a profit on net token sale.

They maybe running at loss after all the salaries and stock comp, but tokens are in profit now.


It's like witnessing a rocket using the most powerful engine on Earth then once it escaped orbit turn off the engine and said "It is flying without power!".

Yes, sure, right now it is ... but that's NOT how it got here.

There are trillions invested to recoup and at most billions in sales. It doesn't add up to tokens making a profit any time soon.


The problem is, people see "they're not profitable once you account for training" and equate that to "AI will go away soon"

But if all the AI companies stopped training new models, they would all instantly become profitable (and stick around)

The thing that makes them unprofitable, is having to compete (which means training models). If / when enough companies exit the market, the cost to compete goes down and you end up in an equilibrium


Sure, but if companies don't exit the market and FOSS alternatives don't end up being unable to get near them in quality, they have to keep spending on training. And conversely, if the market becomes uncompetitive and FOSS sucks, the winners of the AI arms race are very strongly incentivised to stick their prices up anyway...


> if companies don't exit the market and FOSS alternatives don't end up being unable to get near them in quality, they have to keep spending on training

Eh, the AI companies still have lots of datacentres. For the guys who funded with equity, they could collapse down to just running those as utilities. (For the guys who funded with debt, they'd have to restructure.)

From the customer's perspective, this situation shouldn't result in a cost spike. (Consolidation, on the other hand, would. But that's a separate argument from the one the article attemptes to make.)


How often do VC funded unicorns collectively decide to stop scaling up, shut down all their departments targeting growth and reach breakeven point by becoming low margin utilities that will never justify their valuation?


That's all true, but that ends badly for us either way. If there's competition, training must continue, which must eventually be reflected in pricing.

But if there's no more competition, there's no more incentive to keep prices low, which will also be reflected in pricing.


> If / when enough companies exit the market

That will only happen when the bubble bursts and those companies will exit by going bankrupt


> There are trillions invested to recoup and at most billions in sales. It doesn't add up to tokens making a profit any time soon

But this isn't "a ticking time bomb for enterprise." It's an issue for the AI companies' investors.


Good thing the entire nation's economic growth outlook isn't tied to these companies then. For a second I thought we had a potentially dangerous situation on how we misappropriated trillions of capital.


For anyone doubting this, total private investments in the US grew 2% in 2025 relative to the prior year, adjusted for inflation.

But within that big pie, the "IT-related" investments grew 15.7% whereas non-IT actually shrank 2.0%.


Not really, because investors will sooner or later want to see real returns on what they invested. Tokens are suddenly not dirt cheap and enterprises are screwed.

It's like selling dope, once they're addicted, a dealer could turn the screw on them


That's why it's an issue for investors. Their investment may not payout. But the things that were built will still have been built and available to sell for related purposes, the models that were trained will still be trained, and so on.

If things don't end up working out a lot of people have already been (and in the future will be) paid. It's the investors that will lose out, not the subscriber.


When I compare different foundational models on the problems I solve with AI, the differences are not that large to prevent a switch if the price gets too high. I do this like each 6 months, just to assess what is the risk of getting dependent on one provider. It's not yet worring, at least for my use-cases.


Not if they IPO and some other sucker buys the stock.


That's not an excuse to pillage the commons.

Steal from, you know, people who actually work.


Certainly not trillions. The models costing tens of billions to train are a very new development.


OK hundreds of billions with more than ~200B disclosed for OpenAI, more than ~50B for Anthropic and I have no idea how much in terms of infrastructure from Azure, other neoclouds, NVIDIA, etc. It's honestly hard to keep track of "kind of IoUs-ish" from each other but my points is order of magnitude more than few billions than has be recouped so far with tokens and large contracts.


Tokens can be sold at profit, but 70% of compute expenditure goes to R&D and model training[0]. Inference needs to cover all of that as well as being profitable in a vacuum.

[0] https://epoch.ai/data-insights/openai-compute-spend


this will change as inference demand increases (which is happening right now faster than many people expected)


At the same time, the training paradigm being scaled, Reinforcement Learning, is significantly less data-efficient than next-token prediction. You basically need to run an agent for minutes (or longer if you want good long-horizon performance), only to give it a binary pass/fail - one bit of information.

Inference compute is definitely scaling fast, but to scale RL, training and R&D compute also needs to scale hard. I don't think it's obvious that inference will overtake R&D/training, unless there's a reputable source that states that.


do you have some ref?


They aren't being sold at a loss but they aren't being sold at enough to cover the current losses and the costs. The losses are being passed around in some fucked up circular funding mess which will inevitably collapse into a debt crisis at some point.


That isn't enough. Over time the need for growth and increasing profits will squeeze existing margins.


Open source models apply pressures on the low end of the market. The paid models are so much better that they can charge based on value for enterprises.


I wouldn't call Kimi K2.6, GLM5.1, DS4 or newer Qwen models "low end". I prefer GPT5.5, but if it disappeared tomorrow, I'd be perfectly fine with any of these chinese models.


Have you used any of the recent models? My experience with GLM 5.1 does not make me miss Opus at all.


I think for a while this is possible - the models definitely aren't as efficient as they can be as we've seen a lot of promising papers over the last year about how people are changing pieces and parts to do more with less. None of it has come to market yet that I'm aware of so for now it's just a hope I suppose but things like Opus definitely burn a ton of compute to be the leader in benchmarks but the gaps are closing.


In other words, AI companies have positive earnings before expenses


"We'd be making money if we didn't have to manufacture the product", something like that?


He's an interested party. His investments are worth a lot more if he says that tokens are sold at a profit. I don't understand how anyone would trust him?


There are plenty of various providers on OpenRouter serving very large Chinese models like GLM for a fraction of what OpenAI/Anthropic. Presumably they are making a profit.

It’s unlikely that Claude is proportionally that bigger and more expensive to serve so profit margins on inference must be pretty decent


Do we know they are making a profit though? They could be subsidizing use to build market share the same way. They might not have billions, but at the volumes they are selling maybe they’ve got the cash to do it.

Even if they are “profitable” how many Uber drivers are “profitable” because they aren’t correctly calculating asset depreciation. Maybe these guys are doing the same thing.

Maybe it’s a lot of people who already had GPUs for crypto mining, and they’ve moved over to this, so that if they need to grow and buy new GPUs the costs would dramatically grow.


also, it's very much possible that the chinese companies get heavy investments from the state. Since it's very hard to get this info we have no idea wether they really make a profit or not.


I agree, and find that very plausible. I mean, for the CCP a few billions to subsidize domestic AI companies is a tiny investment with a potential huge payoff. It prevents (or at least make it harder for) US companies to build a monopoly on LLM tech and it could help popping the bubble which would weaken the US economy. In fact, if I remember correctly, the AI infrastructure build-out is what is keeping the US from a technical recession.


The R&D is of course subsidized but a lot/most(?) of these inference providers are not Chinese


If the Chinese companies are subsidized anyone who wants to compete with them has to match their price.


> subsidizing use to build market share the same way

To an extent maybe, but that market is almost entirely commoditized already. Besides Cerebras and maybe Groq (which already charge a slight premium) all the other providers are more less interchangeable.

> Maybe it’s a lot of people who already had GPUs for crypto mining

I’m not sure the type of GPUs that were most popular for crypto are at all useful for LLMs?


>interchangeable

If there’s a few providers subsidizing, that’s the price ceiling. Everyone who wants to compete has to subsidize.

Now if this market had been operating for years, I’d say that it’s likely all these companies are profitable or close to it. But the market is so new and there’s so much hype, I find it very plausible that none of these guys are making a profit and they all hope to just hang in until all the subsidies go away.

> I’m not sure the type of GPUs that were most popular for crypto are at all useful for LLMs?

There’s some overlap. I’ve definitely read about people repurposing.


Do you think it will be the case for the Claude Code/Codex tokens as well? I think those are heavily subsidized, but they're the only ones I find real value in.


This is the sort of uncritical thinking that inflates bubbles in the aggregate.


Compared to the inference prices for open models it’s highly unlikely OpenAI/Anthropic are not making decent amounts of money from inference.

How many times bigger could Opus be than GLM or Kimi, it’s certainly not proportional to the price


    it’s highly unlikely OpenAI/Anthropic are not making decent amounts of money from inference.
Based on what? Why are we all whispering about how profitable all this is? It is the absolute last thing these firms would keep secret.


> Why are we all whispering about how profitable all this is?

Nobody is whispering about anything. Everyone is loudly assuming what's convenient for their thesis. Even if you have access to the books, the accounting isn't straightforward–there are yet insufficient data for a meaningful answer.

> It is the absolute last thing these firms would keep secret

If you find an optimisation strategy that you don't think your competitors have, you absolutely keep your margins secret for as long as possible. Knowing something is possible is the first step to making it so.


Based on what I said. If e.g. Sonnet (assuming it’s significantly smaller than Opus) is unprofitable why are there a bunch of inference providers on OpenRouter serving very large models way cheaper? They don’t have a pile of money to burn for no reason.


Brad Gerstner might need a primer in asset depreciation.


Is sounds very much like "trust me bro" ...

Obviously I, like basically everyone else here, don't have access to Open AI or Anthropic books so it's just guessing based on public available evidences, but "tokens aren't being sold at a loss" does not imply there is any profit.

And, even if there is some profit, it needs to be big enough to at least pay back the capex spendings and finance the next model iteration.


so because some1 said something it becomes true ?

/o\


Ignoring the hundreds of billions of investments and debt and the astronomical costs of training and building data centers, sure. This is delusional thinking.


If tokens weren't being sold at a loss, Anthropic would be screaming about it from the rooftops. They've been desparately trying to make themselves not look like a money furnace lately, but it's not really working.

They might be sold at-compute-cost, but that of course ignores training, salaries, and everything else.


Can this project run for 30 years at loss? Google investors don't like that.

One day an exec will say lets reduce wasteful projects and cut this.


You can use Codex and Claude code for most of the tasks that you would manually do

Filing JIRA tickets, updates. Opening PRs, having AI review PRs. This will all use tokens.

No need to tokenmaxx, you will end up burning tokens with just regular AI usage


Not true after Dec-Feb. Opus 4.6 and Codex 5.4 are real.

If you use AI to ship more features, even your competitors are going to use AI to ship more.

So is this how we increase productivity?

Who captures the value? For now looks like the chip makers and model makers.


    > If you use AI to ship more features, even your competitors are going to use AI to ship more.
Unfortunately people didn't learn yet that shipping more features is not equal to increasing revenue which itself is not equal to increasing profits.


But growth stall, because competitors will capture future growth. You need to keep up.

Of course I acknowledge if Hersheys uses AI to make trillions of chocolates a second, there is an upper limit on consumption. And hence a limit on revenue.


> But growth stall, because competitors will capture future growth. You need to keep up.

You need to keep up if you follow Silicon Valley business model growth + funding = IPO. Every other business in the world needs revenue and profit to keep the business open, and growth does not necessarily translates into revenue and profit.


I think social media companies don't need that

Enterprise software companies selling definitely need it. Customers ask was this tested? where is the test report?


It’s not even optional as soon as you’re getting close to any type of standards or compliance framework like SOC2 and the likes.


Management comp is tied to numbers go up

You start slow, then push it the limits

Netflix, never ads to some ads, then eventually its just Adflix, after 20 years.

Each new manager wants that comp up. So ads up by 5% every year.


Lets take the example of Uber. If Uber ships 10x the code and features, I will still not 2x my rides.

Even if Uber makes the cost of travel to 0, I will still not 2x my rides.


Uber needs to prove that they are growing though to validate their stock value, one of the tricks used to be increasing headcount to show growth.

But other tricks include new ventures, essentially public companies and VC companies have an almost unlimited appetite for new ventures, as that is how they keep validating their future growth and stock prices.

Currently financial realities are forcing layoffs, and the AI story is covering for the "growth" validation to keep stock prices going up.

But what's next? After you've fired everyone, what's the next growth story? They'll start hiring again, for new projects, even if AI can handle the coding there is still gobs of work surrounding building a software business or department that needs meat moving it forward.


The end of ZIRP (cheap money) is precisely what ended the new-ventures/new-projects drive among big companies and turned them all to cost-cutting and maintenance mode.


This is precisely my experience. Except our roadmaps didn’t change much.


0, yes you will. Or at least most would unless piblic transit were a genuinely better way to get around. But it won’t be zero as it’s bounded by the base cost of operating the vehicle.


How would they lower the prices to 0 in this scenario, if they have to keep operating their AI-driver-army?


Lowering the cost of travel to 0 would mean implementing a technology by which anyone can simply desire to be somewhere else, and they will instantly teleport to that new location.

returnInfinity is simply lying about not doing double (or more!) the amount of travel in that case.


I’m sure that he/she would ride a lot more if it cost nothing, but I think the point is valid: even if Uber could 10x or 100x productivity, they could not do the same with income, because there is a limit to how much people actually need to go places.


That’s true but fully autonomous driving alone might double my car travel. Going into the nearest major city is a pain. So is driving into the mountains. Operating costs and time are still costs. But not having to drive would really change the game for me.


Who said anything about instantly teleporting? Uber could cut the cost in money to 0 but still operate cars which are bound by the laws of physics and the rules of the road.

Maybe returnInfinity already spends 12 hours a day in Ubers, or otherwise has them satisfy all his transportation needs, and couldn't usefully double his usage of them.


Uber can cut the price of their service to 0.

It's impossible for them to cut the cost to 0 (without using magic), but that doesn't make it impossible for us to talk about what the cost being 0 would involve. Travel time is one of the costs you pay for Uber's service. That you don't pay it to Uber doesn't matter. If Uber reduced that cost to 0, you would use Uber a lot more.


This is a silly easily falsifiable example, though. My Uber usage is approximately zero. If cost was zero my Uber usage would increase.


And look at what that led to. Democracy is not a panacea.


I don't get the atlassian backlash on HN

JIRA is just fine, gets the job done. Its not slow like the on-prem. Jira cloud version is fine.


We are using very different versions of JIRA.

The cloud version I used was slow and riddled with bugs. Entire views sometimes just refused to load or render, or something.

Did it "get the job done?" Yes, in a literal reading of those words, I suppose it did, but anyone who understands the amount of work that a modern 2.4 GHz CPU should be able to do per unit time would not think highly of it.

Nowadays … my company uses Linear … which, while it does have a sleeker, more modern looking UI … is nobody able to make a good bug tracker?


I haven't looked at it recently but always felt that tools like Jira encourage poor management. Effective teams are small teams, and small teams should be communicating and working together across issues freely and frequently. It's generally harmful to have a manager assigning tasks outside of the actual day-to-day discussion without having to speak to someone directly or preferably in an open chat or thread where people can see the discussion.

And ideally the user facing chats and threads are directly linked in to the development chats or at least with channel notifications somewhere.

Task assignment tooling encourages managers to stress developers out with low priority tasks that often start off with incorrect requirements that the structure makes it hard to correct because it is then directly a disagreement with your boss and there is inherently not a discussion it. Whereas a chat at least has the concept of an informational response to a nonsensical task as being fairly standard.


Jira is not a task assignment machine. Jira is, in fact, designed to facilitate exactly what you claim effective teams need to do.


> Its not slow like the on-prem. Jira cloud version is fine.

Seems like opposite land to me. Back in the day running Jira Server was the only way to get a snappy Jira instance. When they discontinued Jira Server to force everyone to the cloud it was god awful slow and forced us to abandon not just Jira but our entire Atlassisn stack.


My company has an on-premise instance of the old Jira and is migrating to Jira Cloud, and the cloud version feels way slower to me.


I don't 'like' Jira, but it gets the job done. It's so easy to onboard users and assign tasks/issues across orgs. Structure is fairly simply and the filters with subscriptions is powerful. Android app that I use on my work phone just works.


Have used both and personally find project tracking thats integrated into the git forge a lot better.

Have used Gitlab+Jira in the past and now only use Gitlab and like it a lot more.


Sergey said that Google was able to rapidly respond to ChatGPT because Sundar had already invested a lot into AI.

They have the people, the models and the chips. And Sundar helped achieve that.

That's what he said on some podcast.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: