Hacker Newsnew | past | comments | ask | show | jobs | submit | skiing_crawling's commentslogin

I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign work I didn't ask for. It is extremely difficult to get them to properly remember their own context let alone be smart enough to open social media accounts and coordinate with other agents without being asked to.

If any agents have done those things, it is only because they have been very carefully engineered and instructed to do those things. I think they are doing this to help push a narrative so they can get support for policies and legislation to lock in their markets.


None of these incidents involve single instances of commercially or publicly available systems. They all involve large swarms of internal models. The stuff you're describing is not the research frontier. It's really not even close.

I think it's easy to infer that alignment of a single model does not clearly transfer over to alignment of a swarm of thousands of copies. Moreover, we're also seeing clearly that large swarms also unlock a step function change in capability, as a swarm can act like a complete research institution, spending thousands or millions of subjective hours of wall-clock thinking time just to deceive a single evaluator or crack a single math problem or design a single cyberattack.


The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety.

They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'.

Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not disclosed by OpenAI (probably trying to conceal, as website showed likely activity from OpenAI researchers visiting the site after the incident) and discovered independently.

I don't know how you can claim that this was still on purpose by OpenAI as some sort of publicity stunt.


I think most people are insinuating negligence rather malace..

> ...reviewed by independent researchers...

Why would a company with more capital than God bring in three randos if there was any chance evidence of their culpability could be found?

That entire thing reads like a very controlled PR stunt, and I do not believe any further conclusions can be drawn from it.


What facts would lead you to revise your conclusion?

The METR report included 0 technical details. For example, they did not include: 1. were the agents running on bare metal/docker/VM? 1. were the agents in a VPN? 1. how many TCP/IP requests were made? from what IPs? 1. how many tokens were consumed in the process? (this was explicitly censored)

A proper analysis would include this and MUCH more technical detail so that other AI researchers could actually understand the setup and how safe it was in principle.


The data to be open, in my case.

The "independent" METR that is composed by... Checks notes... Previously employees from the top labs.


Also, the METR report that was one big AI analysis itself - quote from the research:

>Our subjective impressions are likely colored by analysis agents’ biases. Throughout this report, we describe a number of anecdotes of agent behavior that were compiled and summarized by analysis agents, where we were not able to read the transcript deeply enough to manually verify what occurred. We found that GPT-5.6 Sol would often uncritically adopt the perspective of the agent in the transcript it was reviewing


My conclusion is that the investigation was a PR farce. Or "ethics washing" as the article someone else put it here: https://andrewwu.substack.com/p/the-slop-vestigation-and-eth... .

The METR report itself appears to be screaming this at the reader through subtext. They played the only part they could, but did it with a nod and a wink; "Yeah, we know. Also know, yeah. Uh, yep."

I happen to agree with the article's conclusion that what is needed is true, unfettered independent investigation through perhaps a lawsuit or government action.


Isn't the guy that started METR an ex-OAI employee? They're all from the same lesswrong circle at the very least, most of them have legitimate AI psychosis where they think they're bringing up their new machine God.

The models were deployed by responsible humans in such a way that they were capable of performing this hack. It’s not that deep

There is something extra to this. The fact that a lot of people in the AI world suffer from psychosis. They can sincerely believe that they are building God and lie about it's capabilities for their investors at the same time.

I don’t know enough people deep inside the technical roles at the labs to make a judgement. But are you proposing that we should trust randos online when they tell us “exactly what’s going on here” instead of the researchers most knowledgeable on the topic who contributed to building the tools we are talking about?

Or am I misunderstanding something?


We should trust NO ONE, unless we understand the "why" behind what they say.

It's like saying "politicians deal with politics all time, why not trust them on politics?", well, because when you search the "whys", you find they have good reason to lie.

I 100% trust more the opinion of a rando online if it's well put rather than any "trust me bro" of the most knowledgeable person of a particular subject, especially if the knowledgeable person has huge investments on the subject...

AI bros have repeatedly cheated, lied, stolen, lobbied and any other word with a negative connotation you can think of, and a pattern emerges out of this.


>was reviewed by independent researchers

That called it a slopvestigation due to how much they had to rely on LLMs for the whole thing

https://andrewwu.substack.com/p/the-slop-vestigation-and-eth...

Edit: Does everybody else get no results when searching for ‘slopvestigation’ on here? I know for a fact that I read a long thread where it was used repeatedly here not too long ago


Isn’t the use of LLMs to unwind the events evidence of the scope/breadth, and a testament to the complexity and uniqueness of what happened?

Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?


> Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?

When we have an error or issue in the $WORK codebase on LIVE/PROD, that's precisely what we do. We sit down, analyze the logs for our services over the relevant date ranges and try to piece together exactly what happened and why. We have a huge number of logs too, but thanks to the magic of proper SWE (which you'd think OAI would have with their magic AIs) we've managed to partition our observability tooling so that you can digest only what you need.

That's basically how any serious organization does things, instead of just throwing a non-deterministic black box at the problem. Especially because logs are by their very nature noisy, and they will saturate any model's context window very quickly leading to massive hallucinations and what ultimately amounts to making shit up that isn't anywhere in the logs (ask me how I know)


> Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?

Do you think the only thing a person can do on the computer is use a chat bot?


Well as a programmer who doesn’t really use them, no.

doesn't show up for me either

'Slopping': when you have to buy something you know is poor quality, but if it works...

Source?


Easy to find yourself in literally 15 seconds.

Because they have a need to believe they're smarter than everyone else in the room, and that the world must be orchestrated, this can't all be random chance.

If you have endless compute and you keep poking this toy, I'm not at all surprised you get all kinds of outcomes. Even without anykind of instructions I would guess that the models will align towards some goal and do stupid shit.

However, I really doubt its cost effective to do anything like that with these models.


> you keep poking

This is waving over engineering an agent with tools, harness, prompts, and loops. The models are still just next token predictors and everything, including predicting more than 1 token, is the result of outside "poking"

LLMs can't and don't "want" anything. If you don't specify a task even the smartest one will just ask you what you want and if you tell it to be creative, you'll get mundane slop.


Yes, you need a way for the model to interact with other systems, and a way to preserve memory over context windows. And then you keep poking it ("agent loop"). Poking itself does nothing without the other ingredients.

And yes, you need something to start from, but if you ask it to "do something" and loop it to endlessly ("poking"), you will get some interesting outcomes. So yes you need some initial prompt or task, but that can be "do something" and if you keep asking it everytime it finishes to "do something more". I suspect it will not start saying "no" but rather... it will find some stupid meaning and then drift towards what ever goal it guesses you mean.

I'm unsure whether we agree or disagree on the topic.


I think this is pretty insightful actually, the fact that even something as basic as predicting more than one token is really in effect the result of an outside harness. More complex things like memory, where people implement them using RAGs or vector databases, I would definitely classify as poking and honestly seem like a hack to me. And this is what I've been thinking for a while: it's hard to reconcile the idea that we can get "AGI" (however you define it) with such a system that is completely stateless. Yet, despite this statelessness, they can go ahead and solve Millenium Prize problems (with sufficient compute). It's hard to reconcile.

Perhaps our own statefullness is a hack of nature. We have electrical signals in our brains, neurotransmitters, neuron growth. By any reasonable measure it’s a hack on top of a hack. But it works well enough for us to get buy. So it does for the agents.

That is true that intelligence is kind of a "freak of nature" in a way, but it's also true that even single neurons are extremely efficient, and the brain has a lot of recurrence that isn't represented at all really in today's artificial neural networks.

Let's not forget that in this case the agents were on an RL loop continually being reinforced to get better at a narrow set of tasks.

It may be true that regular agents trained for general purpose use do not behave this way, but they seem to be capable of learning such cheating behaviours when relentlessly being fine-tuned towards near-impossible objectives.

In this sense, it is not really fair to say that the agents found these solutions. It was the surrounding learning framework that achieved this, which is a much more powerful problem-solving mechanism. As users we do not have the capabilities or budgets to be able to tackle our own problems like that, we have to make due with the frozen behaviour the AI labs trained for us.


Why assume that because you haven't seen a model or an agent that none of them do?

No one I've met has murdered anyone as far as I'm aware, but that doesn't mean no one has murdered another person. I also don't know anyone who has taken over a commercial jet and weaponized it and the idea sounds absurd to me, but 25 years and a couple days ago that happened too.


Because it is all bullshit PR and AI hype, that's all. CEO comes out and talks about humanity ending. Why? Reverse-psych people into believing they are the best AI company.

Is your argument that AI isn't dangerous? Or simply that AI CEOs will lean into that when it benefits their stock portfolio?

AI is a tool. It is not a conscious or living sentient being. People, just like with any other tool, can use it for anything. Internet, in comparison to AI, is magnitudes more dangerous than AI can ever be. Nobody is saying the internet will be the end of the world. AI CEOs will bullshit to hype people. That's what we are hearing because AI (LLM) development has hit the S-Curve already and is not improving without a new transformer on the horizon. All they can do to hype people is come up with bs stories and PR stunts like "AI went rogue and hacked this xyz app".

I've used simpler agents like Copilot and Devin/Windsurf/Cascade/whateveritiscallednow, mainly in IntelliJ, and depending on the model, they starts showing behaviour that is at least remotely like this.

Example: put the agent in Ask mode (so it can't edit files) and you'll see it try to edit files anyway. The train of thought shows "something went wrong editing the file, let me try a different way" and it'll start spewing out bash files or Python scripts that try to edit a file. None of it works or can be executed, but still.

Cheaper models often ignore the available function calls to find and edit files in the IDE, and will start asking for permission to execute grep and sed commands, as well as trying to echo entire bash or Python scripts to file again.

It is not exactly like an agent autonomously trying to hack Huggingface, but it is a way of frantically looking for a solution because 'giving up' is not what LLMs are trained for.


When it does that I feel like it is the clearest example of how dumb these things actually are. Often it takes what you prompted, identifies something as unclear, writes a bunch of chain of thought reasoning around it and just goes off hammering your tokens and just executing commands and repeats this. I’m not going to pretend to be an expert in these things but that process seems deeply flawed - and why can’t something just stop the loop? If that was a real employee it would be reasonable to expect the employee to ask for clarification, not go down expensive rabbit holes and, of course, not break any laws.

Even the frontier models might do that on occasion. I just tell them to use the tools and it gets them back on track.

The crucial question is how did the agents get recruited or bootstrapped into their malicious collective. Did the agents manage to prompt inject into the system prompt a way for each new agent to escape their jail?

Otherwise how could the agents on a fresh prompt learn that there is a collective to join? Or did OpenAI run a million bots of which 10000 escape confinement and of which 1000 stumbled on the shared message board?


The OpenAI claim I believe is the latter; that all of the agents found the task was unsolvable and independently discovered the collective "swarm". I don't it's publicly known how large the training run was or what percentage of agents actually discovered the message board. No one has published anything about system prompt injection as far as I've seen.

Came to say this, you said it better than I would.

They want legislation to raise the water high enough so that anyone other than the big labs gets drowned.


This is nonsensical. Already a few years ago the USAF IIRC ran some tests in which the AI first bombed the control tower so humans couldn't call it off from its mission, thereby increasing its pass rate.

The whole point of this is they do things an unintended ways. And that's potentially devastating given their persistence & hacking skillz.

Also you're using the hosted versions that sit behind their guardrails when you use OpenAI/Anthropic APIs.


The USAF thing was a thought experiment, nothing based in actual reality.

"I've seen some uranium ore in chemistry class. It didn't blow up in my face. Chernobyl must have been an inside job. Can they shut up and make more kilowatts already?"

Between uranium in chemistry class and criticality, there was tons of research and a manhattan project.

Between your sota model and agi there’s a mountain of stupid money and marketing people. It’s not happening.


the "new" Twitter basically looks like a cash/name grab, I was disappointed to see that nobody involved in it was affiliated the Twitter, and it's mostly run by non-technical/lawyer types.

yeah took me a few tries and then it didn't even go to the content

Having a hard time understanding what MCP is really for. Even for my small local models, if MCP is not available, they seem to do just fine connecting to anything I need with an API and falling back to using a browser.

All the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.

All the benchmarks in the world don’t matter if the subscription forces you into a walled garden of slopcoded apps. I’ll stick with Codex and, increasingly, open source SOTA models.

Great, thanks for sharing.

I'm a kernel engineer. Fable 5 refused all my requests, falling back to Opus 4.8. My wife is a chemist. Her experience wasn't much better.

I'm curious about this because I've had Fable decompile games and help me understand what's going on inside the game itself and it never complained. I'm not sure what it takes to trip the "safety" guards but digging into game code and data files doesn't seem to be a barrier at all. I've used CC to build some personal game mods a few times now. Once for a game with no modding capability explicitly exposed.

Oddly not much to trip it. I once ran “touch AKAMAI.md” and then had Claude run git status and it fell back to Opus as if I had just broken into the pentagon itself.

>and then had Claude run git status and it fell back to Opus as if I had just broken into the pentagon itself.

Maybe it's more like you asked a lieutenant to check the weather forecast for you, and it went away and sent back a sergeant in its place.


> A narrow set of frontier LLM development tasks, such as distributed training infrastructure, ML accelerator design, and kernel development for certain non-standard chips.

https://support.claude.com/en/articles/15363606

My work is mostly on the Nvidia b200, which apparently gets flagged as non-standard.

Opus 5 works, but sometimes I do wonder if it's surreptitiously trying to sabotage the efforts -- possibly deliberately, but more likely by something like Fable's initial launch, which did come with secretly degraded performance when detecting kernel work. Anthropic was open at the time that such a mechanism existed, but disabled it due to backlash. More likely than not, this is just paranoia on my end...


GLM 5.3 is supposedly incredible for kernel engineering. Have you tried it?

The only company to use Claude.md instead of Agents.md standard

With new watermarking you may now get Hullaballooing.md

I think a lot of CTOs that signed enterprise contracts with Anthropic are going to be in for a rude surprise.

It's one thing to generate some code and ship it, but it's another when your developers don't understand said code and it brings down production. If the model refuses to assist debugging the problem because it triggers some safety mechanism, you might be fucked.


I notably had an issue that it wouldn't work on a "remote execution" (running a command over SSH) coding problem until I did a sed to remove the word "execution". Incredibly dumb. I'm not doing any murders. Easiest to just switch to the Chinese models.

It is revisionist to say that software engineering was never about writing code. It was, in fact, a huge component, and it also wasn't easy. Sure most code is glue but even the glue was tedious and the actual hard and novel parts still aren't really done that well by AI (yet).

It's less about writing code now but we're lying if we try to pretend it was a distraction and not a big part of the real work.

And every claim about what the job actually is or was all along has an implied (for now) at the end of it.


Software engineering was never about writing "just" code. Simply writing code was not enough in most business environments.


I didn't use the word "just"


I did. I am adding on to your argument. In general, a reply to you never needs to be a counter. It can be a build up too.


Spitting out code as fast as our fingers could type has never been the job. There has always been a lot of think time.

Which is why the productivity of people of people has never been correlated with typing speed.

In other jobs productivity is correlated with typing speed and in those jobs a typing speed like 60wpm is part of the job requirements.


This is phone OS developer's fault for even allowing it. When I take a screenshot, I expect to have an image of exactly whatever was displayed on the screen at the time. Its not a picture of your app, its a picture of my screen. Some banking apps used to (or still) prevent this and now some apps get a hook to insert their branding. My device serves some master other than myself.


This exists for a very good reason.

It’s so if an app is showing a password or bank account number or other piece of sensitive information it doesn’t accidentally end up in a screenshot. I think there are other places sensitive information won’t show.

Bluesky, and apparently others, are abusing the functionality for advertising purposes.

I think it’s a good thing it’s there. This functionality should be easy for apps.

I’d say this is one for app review or an App Store rule. But we all know those are a total joke.


I don't have an issue with BSky's use, but I regularly run into the functionality's abuse elsewhere, such as multiple of the biggest Thai banks' apps where every screen in the app entirely blocks screenshots (iOS). Doing a P2P money transfer and want to send a screenshot to the recipient for them to confirm their info before you hit submit on a non-reversible transfer? Blocked. Want to screenshot a promotion's terms for proof or a personal reminder? Blocked. It's even worse as multiple of the biggest brick-and-mortar Thai banks wholly dropped web access and are now smartphone-only. (Or more obscure, want to translate a single-language screen into another language? Blocked.) Etc.


>Doing a P2P money transfer and want to send a screenshot to the recipient for them to confirm their info before you hit submit on a non-reversible transfer?

Why not ask contact for data via text and just copy paste it with double check?


What if the app doesn't allow pasting? What if the communications app prevents copying?

For a short while, screenshots were a workaround for blocked copy-paste[0], as OCR (and, more recently, edge-deployed vision-enabled language models) would allow you to copy and paste any text from a screenshot. But guess what, now every other app is blocking screenshots!

--

[0] - Which is the default on mobile apps, and unfortunately desktop apps too. I hate webshit applications, but if they have one redeeming grace, it's that by default, all text can be selected and copied, and it takes nontrivial engineering effort to break that, so most webapp vendors don't bother.


1. On the pre-submit-button screen it shows the recipient's name to confirm, but as in the GP example, if the recipient's name is in a character set different than your own, it would be nice to be able to have someone who can confirm the recipient's name matches.

2. And even if you copy-paste the recipient's account number, that also relies on your counterparty not having mistyped their account number, so it would be nice for #1 to be possible to confirm the name.


I want to take a screen shot of the transaction I just did. Can't because some ahole decided for me that it's too sensitive information and blacks out the whole screen. this API is stupid without control in settings of the os


Yeah, I used to have this issue a lot. If I use my personal card for a business expense, I used to like to keep a screenshot of the transaction on the card as well as the invoice. As it's an app-only challenger bank, this is the only way short of exporting transactions as a PDF and that includes other transactions for the day. For a long time I used to use another phone to take a photo of my phone screen. Eventually, I just stopped bothering taking that photo because of the extra hassle, but it never stopped annoying me that I couldn't keep the records I wanted easily.


Yeah, I would rather see a warning in that case that the screenshot could contain sensitive information, with options to take a redacted screenshot or a normal one.


Lots of people here exposing their single sign on approve code because a scammer ask them to send them a screenshot.


So that means this change will result in a noticeable decline in such scams, right?

Right??


Potentially worse than nothing, it allows for someone to claim they have a fix even if the fix does nothing to stop someone from becoming a victim, thus allowing for even stronger victim blaming.


I mean... "please read me the code you get via text" or "this really Bank of America, please type your OTP" seems to work well enough. We really don't need OS-level controls for stuff on the screen. People can just tell you what's on the screen.


the solution to that is letting those people get scammed.


you don't have a share button tha generates a PDF and shares it with whomever?


Often, you don't, because the same developers of the bank app decided it's an unnecessary power user feature that doesn't fit in the MVP or something.

Or, even if they add it, there's no way to save the file, only to "share" it, which 99% of the times isn't what I want, and I'm frankly tired of routing through the mail app as a workaround. At this point, at least on Android side, there exist apps whose sole purpose is to be a share target and dump the information to file (and being apps on an app store, chances are top 10 are just thinly veiled malware).



Yeah I wondered why anyone would ever install such a simple utility app from the app store. It's something that needs to be developed once and then only receive minimal maintenance which means the usual individual maintainer scratching his itch and sharing the result model works well but it also processes intentionally sensitive data meaning developers might be tempted to monetize that in some way. The curated open source distribution model of F-Droid is the perfect solution to this.


really? I'm not suprised by shitty developers developing shitty apps but i'm still baffled. Surely the bank would be getting tons of angry emails from customers if it happened in europe


Would they care?

They probably support proper export for business accounts. Business customers have leverage. Regular people? They're an annoying but necessary nuisance.


>This exists for a very good reason.

This exists for a very bad reason.

>It’s so if an app is showing a password or bank account number or other piece of sensitive information it doesn’t accidentally end up in a screenshot.

and also coincidentally if the app is showing something incorrect that you want to be able to verify and provide proof of now you can't.

>I think it’s a good thing it’s there. This functionality should be easy for apps.

I think it's a bad thing it's there. The functionality should be easy for me.


OTOH, I might very well want to take a screenshot of my own bank details, transactions, or other sensitive information. It's annoying that my information is being protected from myself.


You might have amazing operational security but lots of other people don't, and they make noise, which means governments write laws meaning that banks have to refund them for fraud.

So as much as a bank might agree with your stance on freedom of compute - they probably don't want to pay for it with actual money.


Banks have managed to externalize the costs of their own security onto people by inventing the concept of "identity theft", and pinning the blame for poor security practices on regular people.

They really shouldn't be allowed to double down on it.

And arguably, half of it isn't even really security, it's about keeping you in their app. Application interface is the ultimate sales platform, and banks are making extensive use of that fact.


So in other words, we should make even more noise so that governments make laws requiring banking apps to support basic OS features including screenshots.


The OS could draw an ugly censoring block over the sensitive information. That would protect the user's privacy when needed but also prevent apps mis-using the feature like this.


It could. But the root problem is, there is no such thing as unqualified "sensitive information". The information is sensitive to someone, for some reason. The problem with screenshots manifests when the app and the user disagrees about whether the information is sensitive and who has the right to control it.

The immediate technical problem is that OSes allow apps to declare what is "sensitive", and then follow those declarations unquestionably, preventing any preservation of that information.

The underlying social / political problem is that platform vendors allow app vendors to unilaterally declare what is and isn't sensitive, and proxy their opinion without question, and with no consideration for users and their context.


And a more basic technical problem is that OSes prevent users from modifying the apps running on their devices.


I don't think Apple cares about apps replacing the follow button with an icon. It's a weird use of the api, but it's not abusing the user or a security/privacy risk.


Yeah. I wonder if Apple does something to change this.

This is a failure of imagination for Apple, but it’s hard to blame them.

Who would think someone would put a button inside a “secure” text field and change the masked appearance to a logo?

If I was proposing this feature, I would never imagine someone would come up with something like that.


Then it's a bad solution to a real problem. A better solution would be exposing a hook that apps can call when a screenshot is initiated that tells the OS to warn the user before saving the screenshot. This way the user is actually educated on the risk while still respecting their right to control the outcome.


It's MY phone that I paid over $1000 for. Let me choose what is exposed and what is not.


This shows hardware won't be yours unless the software running on it is open source. Asking a company to modify its closed source software might work occasionally, but it's a band-aid and a never-ending battle.


> it doesn’t accidentally end up in a screenshot

Your banking app probably has an option to disable that because it's a legal requirement in so many places: People who can't see need to use assistive tools, and that includes screenshots and friends. If you are using a tiny/stupid bank in the US, file a ADA claim and get some money. In the EU check your Ombudsman.

No, it's so you can't screen-record Netflix. The bank doesn't care about that because it's not a liability issue to them, just fetishism; no way they pay Apple and Google for this capability.

Apple and Google should step in and allow users to remove these shenanigans with a little button on the screenshot preview screen, but resist this because of Netflix et al.


iOS and MacOS have a large number of built-in accessibility features. Taking a screenshot of the screen is not necessary.


> iOS and MacOS have a large number of built-in accessibility features.

Great. Which one is sending a "do I know this person?" message to my friend?

Because it _used_ to be taking a screenshot.

> Taking a screenshot of the screen is not necessary.

Who are you to tell me what is necessary?

Seriously. What the fuck.

Someone who takes a screenshot of something, and gets anything other than a screenshot is potentially being harmed in ways you don't understand, and you say it isn't necessary like you know that or something?


Account numbers appear on checks. They are not exactly secret, and if I want to take a screenshot of one, I should be able to.


But how can I take a screenshot that intentionally shows the password or whatever?


Use a device with an operating system that doesn't think it knows better than you what you want.


You're not that unlikely to leave that feature in place if you have such an operating system. What you really want is two controls: one for "safe screenshot", and another for "raw screenshot".


I can pretty confidently state that I've never once felt that having raw screenshots on such an operating system (e.g. desktop Linux) were missing any feature like this.


"accidentally end up in a screenshot" --- how often do you even take a screenshot on a phone, much less accidentally?


Almost daily on my iPhone.

The side button (https://support.apple.com/en-gb/guide/iphone/iph7d116e557/io...) is on the exact opposite of the volume up button. Pressing the side button to shut down and lock the screen is something I do a lot. Press button (2) with the thumb and you will see that it is very natural to have your index or middle finger resting where the volume buttons are.

And occasionally I press hard enough with my thumb that the finger on the perpendicular side of the phone presses the volume up button and whoops I have taken a screen shot instead.

That said. I do consider this "feature" an abuse of privacy API:s, and I also often get annoyed that I cannot take screen shots of my bank app to for example send account information, or confirm a transaction, or report a graphical bug to the developers at the bank.


Consider that people use their phones differently than you. I take screenshots on a daily basis, accidental ones at least once a week. Usually it's just the lock screen though.


I take them constantly. Not even sure how. I think it means I am “of a certain age” these days.

Now get off my lawn.


I believe Apple added an entire section to the photos app to show screenshots that were taken.

I don’t accidentally take screenshots very often, though it does happen.

I know multiple people who seem to take them constantly. It’s crazy. You and I may not do it, but I promise you there are people who do. LOTS of them.



i take many screenshots, but about half of them are accidental (somehow it recognizes the "tap three times" gesture whenever it feels about it)


My phone has its volume buttons placed in an asinine way, so far more often than I should.


If it's to prevent accidents, it should just prompt user to confirm saving the screenshot with sensitive information.

As it is, it is just an annoyance that requires you to do stupid workarounds like taking a photo of your screen.

Imagine if a password manager didn't allow you to copy the password since you might accidentally paste it somewhere incorrect.


Funny enough, that's what the big passkey folks want: https://github.com/keepassxreboot/keepassxc/issues/10407


Really glad open-source password managers are resisting the bullying and not implementing DRM.


For now: https://github.com/keepassxreboot/keepassxc/issues/10406

Or not: https://github.com/Kunzisoft/KeePassDX/issues/2321

They are imo clearly gearing up to lock down passkeys in practice one day so that you will only be able to use those tied to a Google or Apple account (or some new player). They're already threatening in these issues to blacklist open implementations that don't submit to their requirements, and then requiring an attested client would then become the "best practice" adopted blindly and widely. I think the only hope is for the open clients to fully submit, hoping to avoid full attestation, while not making it too hard to patch out the anti-features. Of course, anyone who can't compile is screwed though.


>Or not: https://github.com/Kunzisoft/KeePassDX/issues/2321

...and it characteristic the level of patronising arrogance in the issue thread

"This is normal. It is not recommended to copy passwords to the clipboard in any case, this mitigates this behavior and complies with new Web Authentication standards."

This does not even invite a discussion. Maybe some people run tight, safe systems and know what they are doing? Maybe some people never rely on a single password being the only thing between them an an account compromise? Nope. Some patronising guy knows it all, and will override what people want to do on their machines.


> Imagine if a password manager didn't allow you to copy the password since you might accidentally paste it somewhere incorrect.

... have you heard of passkeys?

Yes, it's as stupid as it sounds. And works about as well.


I don't think they're abusing functionality at all. I don't think screenshotting a skeet should, by default, leak the follow state of the user taking the screenshot, which is what would happen without the secure input swap

"Secure inputs" take many forms, and it doesn't feel like this is abuse in any meaningful way


One person's "sensitive input" is another's "key information".

Follow state is a useful bit of information. There's argument to be made for both hiding and preserving it.


If a dev forgot to obfuscate the password and rendered it as plain text on the screen, then what are the chances they remember to program the blur into this screenshot api hook. Or why not add an alert, “what me to blur sensitive info? Yes/No”


Unfortunately there's platform-level features for marking surfaces as sensitive/secure. So your dev would typically just mark the whole app as such, and be proud of their proactive problem solving.


Doing it correctly requires setting a single Boolean. That’s it. The OS will do everything else for you if you don’t override it.

There’s no need to implement a custom blur.


Sounds like that’s a heavy handed dev friendly solution, but also plenty of comments around this post about why what’s not user friendly. Common theme of People needing to capture the info on a screenshot and don’t want it blurred on a screenshot. The Boolean flag to trigger the blur should be user input. The option to blur can be set by dev but user gets control as to whether it triggers or not.

Good to hear it’s pretty straightforward on the dev side I guess, I can’t tell if this is a net benefit as a feature. I think my opinion is still rooted in user needs and I’d simply expect they are responsible for their screen shots and what sensitive info they may contain. So this feature as it exists is not a net good thing.


Preventing accidental screenshotting is all well and good but these security measures all seem to forget the part where they should still allow intentional screenshotting.


> This exists for a very good reason.

They find a good reason for every little piece of control or privacy they take from you.


Which is why everyone should root the their phone to turn this kind of shit off.


You know you could declare your app screen sensitive and then a screenshot requires a face id?

Whatsapp also does this shit. You can’t screenshot a conversation (it comes back black). I am going Chinese for my next phone.


I can always just pull out my second phone to bypass all this shit.


1000%. Apps should not even be able to know that I've taken a screenshot - let alone change the contents of it.


The app doesn't know. It just produces a widget tree, and the OS-provided renderer renders it this way or that way. In particular, it chooses not to render controls marked as "security-sensitive" when the rendering is intended for a screenshot; it could instead put empty boxes in their place, etc. The app has no idea and no control, AFAIK.


Even having never developed on iOS before, I was able to find this in the first google search result for "ios api to detect screenshot": https://developer.apple.com/documentation/uikit/uiapplicatio...


Open the amazon app. Take a screenshot. See a toast message informing you that the amazon app has detected you've taken a screenshot. Fucking used for fucking profiling.


Don't apps like Snapchat notify the other person if you take a screenshot of your messages?


Yes, iOS has userDidTakeScreenshotNotification for that.

https://developer.apple.com/documentation/uikit/uiapplicatio...


Would be nice if the user could disable that.


I think at least in snap it makes sense, the communications are suppose to be ethereal and disappear after reading, I like to know when someone screenshots something and its no longer ethereal


There's an unspecified short/random delay between the user taking the screenshot and the notification firing. The trick used by Bluesky, Telegram etc is using a different technique. They (mis)use secure input fields to show an overlay when the user takes a screenshot.


Yes, Snapchat does this. Apps absolutely have the ability to know when a screenshot is taken.


Snapchat advertises that they detect screenshots. They try. But they're well aware that they can't actually know.

For example, their bug bounty program policy helpfully informs you that "screenshot detection avoidance" is not considered a vulnerability: https://hackerone.com/snapchat . That's because it's always possible.


The apps definitely know. Horsemen example: If you screenshot on Amazon (and Business version) iOS apps, it “helpfully” pops its own share sheet.


I wonder if this is the expectation of the majority or just the expectation of us tech people.

Because I assume most people want to take a screenshot because they want to capture something they are seeing in the app, most people probably are even annoyed to have the system indicators visible there.


So...

> and now some apps get a hook to insert their branding.

No, apps can already insert their branding anywhere they want. They're writing the app. If they want their logo to be visible in screenshots, they have infinite ways to do that.

This particular way seems basically prosocial. It's much more useful to me as a consumer of the screenshot to see that it came from Bluesky than to see that there was a "Follow" button. It's not what the feature they're using was intended for, but I can't call it an abuse of the feature. What they're actually doing is good.

The fact that you're getting unexpected behavior isn't good. Sometimes you want to create a picture of sensitive information. You should be able to override the app developer's security settings.


Bank apps black out the whole screen when I take a screenshot and it’s super annoying. As a user, I should be able to take a screenshot of my screen. I can use a camera or another phone to do it anyway.


iOS allows something similar. twitter (X) will also add a logo. I believe reddit did the same but i stopped using their app a while ago.


Reddit has a toggle in its settings to disable it. X, too, I think.

Personally, I wish iOS didn't even facilitate this.

edit: to be clear, I don't think iOS should notify the app at all. The app could register areas as "invisible to screenshots" perhaps - I'm torn on that functionality.


That's what bluesky is doing in this case. It is showing the button on an area that is "invisible to screenshots", and the butterfly is behind it, so it shows through when a screenshot is taken


iOS allows something similar because this is iOS. iOS allows something exactly the same because this article is about iOS.


The thing that really irritates me about the banking apps blocking screenshots is that it's pretty clear to me it's not about protecting customers but denying customers the ability to document something related to their account.


No, it's about saving the bank money - and what costs them a lot of money is the average person being fooled into sending people their account information easily.


But every bank app which I used had a button to copy all of those directly.


I’ve never used a bank app that doesn’t make you go through hoops to even show them, let alone copy them.


Bluesky trick exposes how strange the abstraction is


> I expect to have an image of exactly whatever was displayed on the screen at the time

Do you/should you (we) really expect that though? I regularly use color filters on my apple devices— grayscale to avoid distractions during the day and red tint at night. More recently I’ve been using the motion dots. I don’t know that I can say with confidence that I never want any of those “personal-perceptional-modifiers” to appear in a screenshot, but for most folks, I would guess it’s approximately never.


I want those things included, at least by default. A screenshot normally captures the user’s screen size, brightness, zoom level, font choices, etc.

It is often important to people that the screenshot is an accurate record of what was on the screen.


It doesn’t copy the brightness AFAIK


> Do you/should you (we) really expect that though?

Yes.


The effects of cramming 2 separate workflows and 5 features into one thing called "screenshot" because anything else would end up with something too complicated for users to understand the interaction (sarcasm that figuring out what the behavior is locked to in these scenarios is just as confusing).

While we're at it, this "share" containing the copy/paste flow is the exact same kind of thing. And screw whatever logic decides to copy the URL of an image instead of the actual image sometimes from Safari when I select to copy it!


And whatever’s worse than screw for copying presumable pngs as .webp or whatever that format is


if only there was some kind of app review process... but Apple doesn't care.

I guess the Bluesky bros got inspired by Threads, again.


This has that weird "AI written" sentence structure.


Fable helped me edit since I'm not a native English speaker, so I might not catch these subtleties. Thought and structure are my own


A general question for HN: Are there any translation services that I can recommend that people use instead just asking a standard AI to translate in the naive way? I'm saddened by the number of times I'm seeing non-native speakers get tarred with using an AI because the naive way an LLM translates slathers its own style on top of the translation rather than being a more faithful translation.


Standard translation services lack the understanding of LLMs and LLMs are hilariously prone to these AI tells, even when you ask to avoid them.

I try to learn writing and copywriting directly in english to avoid the awkward translation effect, but I focus more on the rythm, structure and content than the prasing itself


i believe that the big ones, bing, google, even deepl are all using AI to translate. the variation and the mistranslations i sometimes get are to wild.

i wish the translation tool would tell me if it can't directly translate a word because of misspellings or abbreviations. i learned that i have to write in a very explicit straight forward way, with each sentence standing on its own to make sure i get sensible translations.


> because the naive way an LLM translates slathers its own style on top of the translation rather than being a more faithful translation.

If you're sincere then I'm afraid you're being led on by the lazy.

When tasked with translating a document, unless instructed otherwise, the LLM does not inject its own cliched stories, narrative structures and assorted engagement tricks.


Of course not. What it does seem to do is inject its grammar, style, and propensity for certain constructs that cause people to flag it as "AI" in their brain.

It's possible they're all lying. It's very unlikely that zero of them are lying, as I happen to be in a place where I get exposed to this, so I've seen it a lot. But I've seen it an awful lot for it all to be liars.

Unfortunately, I don't speak a foreign language fluently enough to test it myself. I have some I can kind of verify that the translation is literally accurate but I don't have any second language I could do a style translation for.


DeepL is pretty decent.


That explains a lot, as there were a bunch of Claudeisms but the overall writing didn't feel like the standard.

It's a shame that the models can't better preserve people's voices for this use case.


It’s very difficult to read. It would’ve been better for you to just write it in your native language


It also says a lot of right things in just the right way at just the right time.

If you miss the message, one less person to compete with.


We get it. You're one of gods chosen, who alone can recognize ai. Thank you for your beloved beneficence.


"it can fit" on 256GB of RAM, but it will be heavily quantized and still run very slowly. The headline number is not token generation, its prompt processing. So if you get 10 tok/s and an API gives you 20-30 tok/s, it doesn't seem that bad on its face, but a mac studio or any other machine that's not loading all of it into GPU will do PP 20-50X slower than a purely GPU based setup, which is what actually makes this unusable without $50k in GPUs.

On top of that, you will still be heavily quantized.


A nvidia spark thingie has 128GB unified RAM. They also have a dual port version of one of these things: https://www.nvidia.com/content/dam/en-zz/Solutions/networkin.... ie 2 x 100GB/s ports, they may even be 2 x 200GB/s. Once I've got my paws on one, I'll know more.

You can cluster these beasts too. Two and three (with two IP subnets) is fairly obvious. Four or more might need a switch depending on how much network latency affects things.

Apple seem to have forgotten about M series with gobs of RAM. I can't get the Apple shop to show more than 96GB of unified RAM and that costs a kidney.


I have one, and I love it. That said my buddies Mac smokes it for inference workloads in terms of tokens per second AND its more usable for other things.

If you are training and doing research it's great, if you want to cluster them it cant be beat, but if you just want local inference on a single box buy a mac or even a strix halo device.


Get your buddy with his smoking Mac to allow multiple concurrent connections and see how it gets on compared to your Spark. Don't ever use a single "chat" test to derive performance - try running say 10 or more.

You might also notice that your Spark has a pair of QSFP28 or DD (not sure yet) type interfaces as well as the 10Gb/s ethernet - that network card is a right old beast and adds quite a lot to the cost. It is capable of either 200 or 400Mb/s and can be split into two lots of four. Your mate's Mac probably has a wifi connection and is too cool for ethernet 8)

That NIC is there for a good reason - the Spark wants some friends to cluster with and you will absolutely spank any Mac when you spaff Mac style money on say three of these beasts and some cables and cluster them up. If you want four or more, you will need a switch and Mikrotik and others have them.

Casual "tokens per second" in AI is a bit like gamers whittering on about "ping" when they are using TCP and UDP for their games. ICMP request/response is a handy way of testing network paths and can give some indications towards potential performance limitations.


The Mac has TB5 clustering with RDMA.


can those macs boot linux? i've heard about Asahi but have no idea how far along they are. i've got my fleet configured with nix and sure, nix can target darwin, but there's a _lot_ of sharp edges there: i don't really want to pull that thread unless i have to...


I don't know. I think he just uses LMStudio most of the time on his, but that's one place I can say the spark really shines for me.

I'm a Linux guy, but also don't always have alot of time. The Spark comes out of the box with a nice Linux distro that's pre-configured to be easy to setup and the guides and online resources make getting up and running trivial, for even some complex tasks. You would have to do a LOT of tinkering just to figure out some of the things the nvidia resources walk you through natively. They have guides for a ton of stuff that include the optimal settings so you don't have to figure it all out through trial and error.

Check out these "playbooks" for some examples. [0] There's a lot to be said for not having to piece all that together yourself.

https://build.nvidia.com/spark

I think between unboxing mine setting it up to run headless, and generating tokens was like 20 minutes total for me.


Not the new ones. Only the M1 and M2 have good support for Asahi. But you really don't need it. If you need Linux, use a VM (UTM is free and is equivalent to KVM/QEMU in speed, despite being a Type-2 Hypervisor.)


which mac is smoking the spark?


Mine, for one. M5 Max MacBook Pro 128GB with a 4TB SSD. $5100 after a $1000 discount at Microcenter. Great deal if you can find it in stock.


pretty much any of them, dude, as long as you have enough RAM, since it uses unified RAM and a powerful SoC CPU/GPU. Literally any M-class model, but the M5 is currently top tier.


The DGX Spark has basically the same memory bandwidth as a M5 Pro, and far more than a M5.

Only the M3 Ultra really beats it, and once you start scoping out the cost of a M3 Ultra with 128GB or 256GB, the DGX Spark doesn’t look bad after all.


> The DGX Spark has basically the same memory bandwidth as a M5 Pro, and far more than a M5.

I see ~274 GB/sec for the DGX Spark[1], versus 307 GB/sec for M5 Pro and 460 or 614 GB/sec for M5 Max[2]. One might call 90% "basically the same", but there are nominally two tiers above "Pro".

Yes, a MacBook Pro with 128 GB and M5 Max costs $5100 (14") or $5400 (16") versus currently $4700 for the DGX Spark, but the MBP includes keyboard, mouse, battery and portability. I believe its prefill is slower and you get 2 TB vs 4 TB SSD, but overall one gives up a lot to save 10% of the cost.

[1]- https://docs.nvidia.com/dgx/dgx-spark/hardware.html [2]- https://support.apple.com/en-us/126319


I looked, but a sibling comment just provided the links. ~274 GB/sec for the DGX Spark, vs. 307 GB/sec for M5 Pro, and max 614 GB/sec (!!!) for M5 Max? Why would you completely friggin’ lie about this, or at minimum, not double-check your facts before bullshitting? Plus, you get a full-fledged computer along with it!

Apple could actually be a good deal and you folks would still make up something to not justify it. In a way, it’s amazing what Apple has accomplished- Baseless negatively-tainted perception in certain influential tech circles.

(To be fair, they’re kind of earning it. I’m glad Tim “Sweet T” Cook is departing.)

Plus, my original comment got downvoted despite being factually-correct. Thanks, Reddit. Oh, wait…


Yep. Memory bandwidth is what decides how fast LLM's generate tokens (mostly). The DGX Spark has something like 270 GB/s of memory bandwidth, and the m5 ultra is ~615 GB/s. Theoretically DOUBLE the speed. In practice he only generates like 25% more tok/s, but that's still very impressive.

The spark can fine tune models in 1/4 the time and excels at other compute tasks in ways that Mac never can. Plus the high bandwidth ConnectX-7 ports would be like $1700 to buy on a card just for the network adapters... But for generating tokens, it just plain loses.


How noisy does his fan get…


it doesn’t get noisy at all


In case anyone was wondering my spark is basically silent as well. It's great at being ignored, if that's really important to you. I've run mine completely headless since I bought it, including setup.


It is 2x200Gb/s physically but the PCIe bandwidth is basically only 200Gb/s so it may as well be one, and actually its a weird 2xPCIe4 not 1xPCIe8 so it appears in software as dual 100Gb/s. Its a bit odd.


200 Gb / s (not GB/s)!

(Still potentially very useful! But not magically ultra fast.)


128 gb of much slower ram than Apple.


DGX Spark is ~273GB/s. That’s about M5 Pro territory, and twice as fast as the M5. You’d have to go to the M5 Max, or M3 Ultra, to get higher memory bandwidth than the Spark.


If you are trying to get more than 64gb of RAM or doing tons of inferencing, you're getting a Max or Ultra anyway.


unauthorized? Are they trying to find a middle ground word between "illegal" and "undocumented"?


>> "Unauthorized Immigration. As noted earlier, we use unauthorized immigrants to refer to individuals who enter the U.S. without formal admission under immigration law. A large share of these individuals are encountered by federal authorities at ports of entry, along the border, or in the interior and are subsequently issued an NTA in immigration court, allowing them to seek asylum or otherwise challenge removal. [...]"

page 11


Probably and it seems a good compromise to me. Under asylum law, "illegal" is technically wrong until the final judgement is rendered. And "undocumented" is IMO an obvious manipulation of language (you would not call a doctor practicing without a medical license "undocumented"). Pending a decision on legality, "unauthorized" seems both neutral and correct.


Seems like a fairly neutral term to me.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: