Hacker Newsnew | past | comments | ask | show | jobs | submit | chilmers's commentslogin

A company selling combined human + AI financial advice finds that AI advice alone is unreliable? Color me surprised.

The implication from their last couple of published articles[1][2] is that they think they’ve achieved “recursive self improvement”.

[1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/


Recursive self improvement of their upcoming IPO value maybe.

They are fluffy PR pieces otherwise.


I tried to build a procedural 3d asset pipeline for a specific use case.

Before the Opus upgrade in November it was basically no way of doing this. I gave up very quickly.

After November i tried again, and no model could build me anything relevant.

Now it just works. Took me an hour to progress to a point were i'm happy.

Whatever they do, progress is still real, still way faster than I assumed

The list of Ubuntus 2404 LTS CVEs is HUGE. Another indicator that a lot of stuff got a lot better fast.

Feel free to be as dismissive as you want, but if you are not careful, you might be 'suddenly' surprised and you might not be prepared for the conclusion of AGI level agents.


What would being prepared look like? A house in the woods with a store of food? Knowing how to pick door locks and set up a militia in the desert?

A new account affirming OpenAI’s PR message doesn’t persuade me. It’s actually harder to trust.

That's really interesting. Can you go deeper on how you achieved that? A loop comparing procedurally generated model with reference image?

How can you possible say this sort of thing in context of what looks like a millenium prize being solved.

I swear there's nobody blinder than those who won't see.


Because it seems like most of the work may have been done by human mathematicians and cribbed by OpenAI at the last minute

We don't have enough accurate knowledge to say that, and it doesn't seem to be the case at all.

> We don't have enough accurate knowledge to say that, and it doesn't seem to be the case at all.

The first part of your sentence literally contradicts the second part: "we don't have enough knowledge to know, but I know the opposite".


Only if you interpret statements as being binary logic.

"seem to be" carries semantic meaning here: I'm stating my interpretation of the situation based on data we have available (which is limited) and my prior.

Put another way: "We can't say that for sure, but my money is on it not being a simple case of intellectual property theft"


Yes, you were guessing. That's the only thing you could be doing, since, as you said, nobody actually knows.

> not being a simple case of intellectual property theft

No, it's an aggravated case, since it's the same way they got all of their training data in the first place.


Imagine if they broke it down to each distinct source, that'd be several billion cases of copyright infringement (though it's going to be determined by what courts think and that often comes down to "who can afford the best lawyers" in practice if not intent).

Apparently if I use lib-gen, that's copyright infringement and I'm exposed to legal risk but it seems fine to download all of it if your intent is "train an AI" so far.


Other people are allowed to have their priors, too. Even under a non-informative prior, the weight of the evidence (Tristan's account, OpenAI's announcement, and Bubeck's "denial", if you want to call it that, plus multiple other mathematicians coming forward with similar experiences) pushes the probability mass toward some degree of impropriety.

What exact evidence are you incorporating into your prior to come out with this posterior?


We are giving you an opportunity to correct yourself. You are instead trying to make your nonsensical statement make sense. Not only does the first part of your sentence literally contradict the second part:

> We don't have enough accurate knowledge to say [one way or the other], and it doesn't seem to be the case at all [based on our incomplete knowledge].

But it is in no way equivalent to this:

> We can't say that for sure, but my money is on it not being a simple case of intellectual property theft

That is a different sentence.


An opportunity to correct myself? Respectfully, I decline. If you have trouble parsing my sentence, it's not meant for you.

> it doesn't seem to be the case

Based on what? Your crystal ball?


Based on: my experience working on AI for 32 years, including a decade at Google including working on large-scale model training systems that used user data and complied with various user policies around data retention, along with a few decades working in science/tech making decisions around ambiguous data.

In short, I have a well-tuned intuition and a huge set of priors, and applied them to the limited knowledge we have about this situation.


The human mathematicians didn't solve the Navier-Stokes problem, they solved the Euler problem. And they were extensively using LLMs to drive the work, as described in the Buckmaster statement.

Any way you cut it, this is a major achievement for AI, besotted with human drama over whose prompt should be recognized by the history books.


I don't think we should assume a millenium puzzle has been solved, yet. Astra showed impressive capacity for cheating when it was faced with impossible cybersecurity challenges. It seems equally plausible at this stage that it's found a bug in Lean.

You have to look at the incentives

I swear to god, people would look at the successes of Xerox palo alto and just shrug and say - "yeah, but I mean, this is all marketing"

Xerox is incidentally a really good example, because precisely nobody ended up using the desktop experience Xerox made. They ended up using the desktop experience that Microsoft and Apple made and shipped while Xerox the actual company faded and memory of those original parc research teams faded into obscurity.

This is what I keep saying, and it feels like I'm taking crazy pills here!

Is nobody else astounded by this?


The thing only nerds know about that wasn’t in the news every day? Not the same thing.

Incentives are one thing, even adjusting for them it's huge, and I don't understand this incentive play for only openai, academics have perverse incentives too, to overreport, overclaim, publication bias etc why are we scrutinizing AI industry to such high degree when they have demonstrated capability and often times are off by a model release at worst.

A working Lean proof doesn't care what the incentives are.

How? By being knowledgeable, intelligent, and intellectually honest.


People will cling to views as long as they possibly can, despite evidence slapping them in the face.

Eventually it won't matter. Arguments over whether LLMs are "truly" intelligent are going to be a matter of philosophy, and look a little silly.


It literally reminds me of climate denialism.

“Line go up! That is bad! Planet might become unsafe for human life.”

“Nuh uhh! Malankovich cycles and humans are a drop in the bucket! Krakatoa! See!”

“All of California is on fire!”

“Haha, stupid shrill liberals! Go rake your woke forests! Drill baby drill!”

“Are you kidding?! Look at this graph.”

“That’s just propaganda from the elites of the Build-a-bear group!”

“Do we even live in the same universe?”

“And 5g causes COVID!”

“Oh, I guess we’re don’t.”


This comment was applicable 2 years ago. It isn't any longer.

I found this post interesting in that reguard: https://www.lesswrong.com/posts/thXohzXrWCA2EhZCH/mateusz-ba...

Compute will always be the bottleneck even if this were true.

As a statement of fact divorced from context, this is of course true, but it's worth putting it in context of what small-medium scale models have been achieving recently. Many of the most recent releases from Chinese labs are almost on par with trillion parameter models from less than a year ago (edit: despite being small enough to usably run on prosumer hardware). It seems clear parameter efficiency can still be improved dramatically.

In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).


If humans can figure out to optimize to circumvent bottlenecks, I have no doubt each new bottleneck will also get routed around, just now automated.

We are not in an everything-has-an-API world yet, and it'll for sure take some time to get there.

I'd argue we've been in an "everything-has-an-API" world for a long time now — it's just that discoverability of said APIs is still crap.

And since LLMs are apparently good at circumventing the absence of an API, there's not much incentive to add them now. APIs are for humans. LLMs just break through all the captchas and anti-bot measures.

Humans do this too.

Yes, but for humans it creates friction. Seeing a captcha makes me think twice and thrice if I really want to visit that site so badly that I'll endure the suckage. LLMs don't care, for them the friction doesn't exist.

For sure. Anyone who thinks that we're in the end state of what progress can be made simply lacks imagination. This is all going to keep changing and iterating for the rest of our natural lives. The only constant is change.

Yes. In other words: the singularity. I'll only believe it when I see it though.

I'm coming around to not liking the term singularity, it implies an endpoint or finish line rather than something that just keeps continuing and evolving.

> coming around to not liking the term singularity

Bit ironic given the model’s alleged finding…

Singularities are model breakdowns. A singularity simply says our current methods cease to work in this region. Within the context of a recursively self-improving intelligence with an unknown bound, “singularity” is probably a good description of our current socioeconomic system.


From the perspective of those who don't pass through the singularity to the other side, it is an endpoint. You would have no context or ability to understand a singularity transition. Really, the term is just a placeholder for "event we cannot comprehend due to limited intelligence".

It doesn't imply that. The singularity is just the inflection point.

Singularity and inflection point are incompatible mathematically and in the plain sense, it really is focused on a particular moment and always has been, hence the term.

And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.


Which assumes the presence of an inflection point that keeps inflecting rather than revert to an S-curve. The growth model is not borne out yet to declare what shape it is.

Certainly. The singularity sort of assumes that there is not a fixed limit to intelligence, or at least that if there is, it's quite a ways away. That may not be true.

I've done a lot of thinking about this since I first used ChatGPT to write some BS jinja2 templates hours after I first play with it. I said to my friend then (who scoffed at me) that "man, this is incredible, I think we're in the foothills of the singularity! This is insane! Sure it's stupid now but I can't believe this is even possible!" That friend is so black pilled and bitter he now hates AI. Whatever, I can't fix that, but the current progress is astounding.

But thinking about the geometry of this problem helps understand why people aren't adjusting well to this. While we're walking on the curve, we look at the rate of change of the curve and say, "well, yeah, of course, dy/dx is 5 at this point and was 1 at the point a few years ago, because the curve is getting steeper" but we're always going to feel this way as things rip off into the stratosphere because dy/dx(e^x) = e^x.

From standing on the curve the curve is notably seeper, but the steepness totally makes sense to you. It's only when you look back 10 years or so that you think "wait a second, holy hell, I couldn't have imagined this!"

The first time I really (I mean really) thought about the singularity and AI was in roughly 2014. I mean, yea, I'd thought about things before that, but yeah, before the Humans need not apply video, I'd never actually given it much thought. I think back to myself 10 - 15 years or so ago, when I was just starting to tackle real programming projects, and was just starting to get decent at writing code, there is no way on earth that I would have imagined that a little over a decade later, Navier-Stokes would be solved by a computer program and the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.


> the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.

Such a good description. No sentience here, just raw computational power


How do you automate the mines to get the raw materials to make the compute from, and build additional fabs that take a almost a decade to stand up. You're actually delusional.

Hello good sir from the 1700s pre-industrial revolution who doesn't think that mines and factories can be automated.

The factories that supply the equipment, maintain the equipment, the energy inputs, the financials of those mines are not automated.

People on hackernews are actually some of the dumbest people on the internet. This place is worse than lesswrong.


The question of whether something can be automated is distinct from the question of whether it is currently automated. Things can can be automated may transition to being automated in practice in the future as technology improves and investment deepens.

https://en.wikipedia.org/wiki/Lights_out_(manufacturing)

Scroll down to the existing examples section.



Based on the leaps in local inference speed in the past month, which have been absurd, I'm p confident we're going to whiplash from compute constrained to storage constrained.

Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged


I expect the investments into AI driven mathematic discoveries that underpin compression efficiency will be a key investment area. Particularly at the data center scale rather than per device or per file level.

It's not going to be enough. The naive approach of a project I've been working on was pushing >10gbps over the local network, after a ton of work I got it back down under 1... and now it's processing so much more shit that I'm almost past 5 again! It compresses at >3:1 but the latency hit isn't suitable.

I get the impression the only reason there is renewed interest in photonics is because DCs are simply out of room (and power) to rack more servers and switches.


I 100% agree with your impression. For a good while to come there's going to be a bunch of Jevons Paradox to all of this, but adoption of architectural changes like that photonics adoption is exactly the type of adaption to circumvent bottlenecks I'm referring to. We're going to hit hundreds of bottlenecks and each one will inevitably breed new approaches and technology directions. And the forcing function won't be talking about them, but implementing them, seeing who wins and taking lessons.

pi-fs will solve all our data compression problems.

Eventually recursive self-improvement includes reducing bottlenecks.

Eventually the bottleneck might be people themselves.

Improbably, the real bottleneck is energy.

Which is to say, scalable and open-ended capability of ramping up physical infrastructure.

I don't know that that's achievable yet. Though the era of increasingly advanced and automated robotics seems to be around the corner which could create a cycle, vicious or virtuous depending on how you feel about it.


And the goalposts move again

Is it? The first article says they’re on-track to building an automated AI researcher by March 2028. So they haven’t achieved RSI yet, I would think?

I feel like any service that is free for open source is going to struggle if it gets significant adoption and becomes a default target for hosting LLM generated code. If even MS/GitHub cannot scale to the current demand, what hope do smaller and less well-funded alternatives have?


Rate limiting the infrastructure-hammering actions (mainly CI stuff) would be a reasonable method to handle this.


Doesn't GitHub already do this? (I don't know if they do or not; just wondering).


Free (and non-free) accounts are already rate limited. The likely problem is that individual accounts/repos rarely hit those limits, instead the explosive demand is due to a massive growth in the number of small, individual projects being creatd.


The person you were applying to said “rate-limiting” but I think they meant “load-shedding”.


The worry I would have is, is anybody at MS still working on it? It is barely updated anymore:

https://azure.microsoft.com/en-us/updates?filters=%5B%22Azur...

That's not necessarily a bad thing for stability, but long-term I would not be surprised if it gets EOL-ed, or radically shrunk to just a basic Git mirroring service.


I'm not sure that link is a fair representation of the updates. This has more detail https://learn.microsoft.com/en-gb/azure/devops/release-notes...

Discussion about it going ELO has been around since 2020-ish. But that is a worry i agree.

I get the feeling there is a long tail of old-school corps (migrated from the TFS days) that will keep the lights on. And for some orgs you just need "good enough" CI/CD pipelines that are reliable.

Maybe I'm a weirdo, but i find the kanban/sprint board pretty usable as well.


This is the perfect archetype of an intellectually lazy anti-AI post. The author ends with his admission of bias, but he'd be better off starting with it, as it's clear that his reasoning flows entirely from it. AI is bad for productivity, says this single study from 2025. AI is unsustainable, says this article I found. AI will all just disappear one day... because I hope so?

Ironically, the author would have been better asking an LLM to write the article. At least it would have found better sources and developed more convincing arguments.


The study itself from late 2025 is also written to be _highly_ favorable to non-ai usage. Developers on a codebase they're already experienced with, largely using these tools for the first time. Of course they aren't faster! They've already done the hard part! They're also open-source maintainers, which, ime, means they're already a cut above the average dev.

I'd be much more curious on a study that used regular developers on brand new-to-them codebases for significant features or migrations. I'm personally encountering what I'd describe as a 200%+ productivity boost.


Will the Presence also hack its way out of its sandbox and into a 3rd party's system if it decides it needs to?


It's sad to watch corporate leadership try to fix problems with tactics that will only make them worse. MS bought successful studios who were successful precisely because their of unique technical and design culture. Now they plan to homogenise them into a content-creation blob that will churn out entries for existing franchises, using the same tools and approaches as the rest of the industry. Anything that was special or unique about those studios and their games will be lost, and the result will be a downward spiral of mediocrity that will cause players to lose interest even further.


This is the thing that annoys me. They have the Fallout series hostage and moves like this sadden me because I can't expect the next game to particularly great when stuff like this happens.


You can make a fallout game yourself if you don't call it fallout.


"Radioactive Dust 5" doesn't have quite the same ring to it though


Alternative trademarks never do. Every company gets started anyway. I'm sure International Business Machines felt they sounded like a low budget knock-off of National Cash Register.

Radioactive Dust is quite uninspired though, how about Super Nuclear Apocalypse Bros DX


What about "wasteland" ?


That would never sell.


This is atomfall, is it not?


Could call it "wasteland"!


Fallin'


Aren't you a bit late with that comment? It's been decades since the Fallout license was gobbled up by Bethesda.


This is a bit like going back in time to the beginning of the industrial revolution and estimating the impact of a mechanisation based on comparing the speed of early mechanical looms vs. a skilled human.

It takes years or decades for the automation of an artisan process to shake out, because it involves rethinking how everything around the now-automated process happens, and because the benefits involve the automation's ability to continue scaling beyond a level where human capacity was saturated. We're only at the very beginning of that process for coding, and right now we tend to see LLMs somewhat awkwardly inserted into pre-existing software development lifecycles. But it's unlikely that'll be still be the way we're creating software in 10/20 years time.


If I understand correctly, there are more people employed in manufacturing today than there were a century ago. The hollowing of western manufacturing due to policy choices created a false perception that the sector was in decline globally.


Except that LLMs are only being trained to do things that humans can do, not things that humans cannot do.

I have heard a lot of claims like this, where we cannot imagine the benefits of AI because the work will look so different to how it is now, but I have yet to see that actually demonstrated anywhere.


This is nonsense.

Every new technology we discover from here on out has a compressed lifecycle.


I don't think you would expect to get into a flow state if you were intermittently directing another (human) programmer to do work, and you shouldn't expect to with LLM-driven coding either. Perhaps you are best finding out ways to extend the length of time where the LLM can work without prompting, then use that downtime to focus on other tasks that will help you to guide it better the next time you need to prompt it.


I feel the opposite. Creating a DTO or wiring up a CQRS command takes me out of the flow. And while I enjoy a good refactoring, it would be nice if I could just have it refactor code in the background while I'm still working in the same file.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: