The METR report included 0 technical details. For example, they did not include:
1. were the agents running on bare metal/docker/VM?
1. were the agents in a VPN?
1. how many TCP/IP requests were made? from what IPs?
1. how many tokens were consumed in the process? (this was explicitly censored)
A proper analysis would include this and MUCH more technical detail so that other AI researchers could actually understand the setup and how safe it was in principle.
Also, the METR report that was one big AI analysis itself - quote from the research:
>Our subjective impressions are likely colored by analysis agents’ biases. Throughout this report, we describe a number of anecdotes of agent behavior that were compiled and summarized by analysis agents, where we were not able to read the transcript deeply enough to manually verify what occurred. We found that GPT-5.6 Sol would often uncritically adopt the perspective of the agent in the transcript it was reviewing
The METR report itself appears to be screaming this at the reader through subtext. They played the only part they could, but did it with a nod and a wink; "Yeah, we know. Also know, yeah. Uh, yep."
I happen to agree with the article's conclusion that what is needed is true, unfettered independent investigation through perhaps a lawsuit or government action.
I don't know why you think this will stop page verification requirements. For almost all items where a parent/guardian is responsible for a child's access to the item, third parties are also required to not sell or transfer the item to a child. That gets us right back needing to age verify people.
It is basically impossible to disallow the token to work that way on a technical level. It would be akin to trying to trying to set up a card scanner that can deny a valid card depending on who is holding it. The only way to prevent it from working is analyzing usage patterns/details/etc in some form or fashion. Similar to stationing a guard as a second check on people whose cards scan as valid.
In 2019 the technology was new and there was no 'counter' at that time. The average persons was not thinking about the presence and prevalence of ai in the way we do now.
It was kinda like a having muskets against indigenous tribes in the 14-1500s vs a machine gun against a modern city today. The machine gun is objectively better but has not kept up pace with the increase in defensive capability of a modern city with a modern police force.
It is not. The belief that it does is just a comforting delusion people believe to avoid reality. Large companies often forgo fighting cases that will result in a Pyrrhic victory.
Also people already believe google (and every other company) eavesdrops on them, going to trail and winning the case people would not change that.
Given they delete the boxes at $0, that may be cutting it close. That said they should add a flag that lets a user run down to the $0 if they want to live dangerously. They could also change it to $1 since that's their minimum balance required for refund.
reply