Ignoring the checkbox is an utterly offensive move, but repeatedly manipulating it contrary to stated consumer intent is a whole other level. I didn't know we were supposed to take 'frontier' literally in every sense of the word.
I have now witnessed this myself after not believing this at first. Of course, screenshots etc. will hardly prove anything. This needs a proper third-party audit!
Reverting... to on? I've seen some stuff in Apple-land where I'm pretty sure it went back to the default option after an update but I've never seen anything get granted more permissions.
No, like specific apps that have been granted full disk access etc lose their permissions. It’s hell for IT, constantly breaks AV. I would not characterize that as reverting to a default. More often it’s a breaking change for how the security model works that ends up requiring reauthorization
Ah, sure, it's annoying but it's never giving new permissions you didn't give it before. And yes "revert to default" is probably not what's happening, but reverting to forcing a specific permission grant when the existing one was superceeded or replaced.
I understand the rationale. The app you granted full disk access three months ago might have significantly changed. You can't reset the permissions for a given app after every update, as you might end up prompting the user three times per week.
The alternative would be keeping a permission granted for an app that might have had an update introducing nefarious behavior, whether intentional or not.
Linkedin, Amazon and others have gotten around this by adding a "new" feature, default on. I've gone into privacy settings where everything is unchecked, except a newly added checkbox.
This is apparently enough to hold up the broadly-accepted fiction that its possible to use these services while maintaining privacy and without risk. There's a huge industry of hosting services to handle the needs companies who compete with the primary cloud providers, who can't afford the IP and competitive risk.
OpenAI was caught directly stealing from apple, asking employees to bring in their laptops. Its a polite fiction that companies aren't trying to gain any advantage over the other. There's no effective consequence, and even if they get caught red handed they can litigate for decades.
You know what they'll say in their defense. "This is an extremely complicated systems, and we apologize that a technical solution was broken in an intricate way. [Insert boilerplate about taking privacy seriously here]"
Pirating copyrighted works is absolutely illegal in most jurisdictions, but pirating every single copyrighted work in the world is somehow exempt from law.
We can't apply plebeian laws or ethics to our benevolent overlords, they are above our worldly worries.
Wall, you can't tell if they're ignoring it, but it's fairly easy to show they're manipulating it. So a cut-and-dry demonstration is worse than a strong suspicion.
I've been suspecting over the last couple of years of the frontier companies using data for training anyway, regardless of training-use consent. "Using" the data doesn't have to mean they literally upload chat transcripts into pretraining datasets. My analogy has been money laundering -- if that can happen at massive scales, surely these companies can and will do the digital/data equivalent derivations/transformations. Even if one could have the access etc. to do so, how exactly would one prove that a given synthetic dataset that OAI/Anthropic uses is derived from particular user conversations that did not consent for the info to be used in training?
Consider, for instance that OpenAI's (consumer) terms say "If you do not want us to use your Content to train our models, you can opt out by following the instructions in this article ." but they also do say "We may use Content to provide, maintain, develop, and improve our Services". [1]
If you think that's quibbling, consider that OpenAI's business terms, in contrast, do state "OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.". [2]
And humanities have a word for this, exploitation, or appropriation, maybe it's time scientists and engineers revisited basic ethical notions. Skimming a dozen threads and nobody seems to have this vocabulary or willing to say it.
There was also this follow-up [1], pointing out that most of these 7 studies only had something like a 50–65% estimated chance of producing a significant result, yet all 7 did.
In other words, if a whole series of fairly noisy experiments all comes out squeaky clean, you have to wonder whether you're actually seeing the whole series.
Yes, evidence is needed but especially for the claim that distillation is what makes these open models good. Serious citation needed.
Think about it: even if they distill the shit out of frontier models, the model still gotta learn, right?
If anything, as you can see from the K2 Horizon release, aggressive (self-proclaimed) reliance on distillation does not result in a model that has remotely any frontier capability. Try asking K2 Horizon to write iambic pentameter for instance, or even give it the car wash prompt. I tried both these on the Q8 quant for the 7B model and the results were depressing.
To whoever downvoted this, it would be helpful if you actually reply with something substantive. The post I'm responding to makes sweeping characterizations, and I challenge it with a relatively good heuristic and indirect evidence, and only get downvoted?
Because it's evident you didn't engage with the last sentence I wrote properly. "Citation needed" in more words is not a sufficient response.
> As for people asking where is the evidence for half of this, you will never have public evidence for most of this, but deduce what the partly hidden parts reveal about the whole and it is fairly obvious, especially if you look at the past behaviors of the governments and other actors.
To help you further understand, the Chinese labs will never admit they were distilling until they are better than US labs with their own non-distilling process or some sort of espionage-like public reveal shows it. So you need to look at secondary indicators. Much like how the chinese government lies about their economic stats so 3rd parties use secondary indicators to figure it out.
My point is simply this: Chinese labs going to great lengths to distill Claude is evidence that distillation is useful but not that that is what makes them so good. That one doesn't even need them to admit it.
Nathan Lambert has made this distinction explicitly: he thinks Chinese labs likely innovate heavily on distillation, while saying he wouldn't call it a crucial factor in their post-training capabilities, in part because the RL work still has to be done by the lab itself -- just like what I said above.
Sounds obvious but just try using the models for anything outside the evals. Take something arcane from Greek history, use it to create a masked linguistic puzzle, which you then ask the model to solve mathematically, all wrapped as an ask to generate ASCII art. Yes, all these elements exist in some form in the evals but the key is in how utterly unconventional the elements are that you pick and in how you combine them.
I have consistently noticed Opus 4.8 and GPT-5.6 far outshine the Chinese models. Gemini is sort of middle of the road, Grok is better than Gemini but not really close to Opus/GPT. OAI & Anthropic still remain unbeaten by a wide margin in my eyes.
Certainly, real-life is the ultimate benchmark. But for various reasons that isn't always immediately possible to go evaluate a model on.
Maybe my Greek idea sounded too high falutin' or simply seemingly clever (I give an example below -- try it out!).
So here's something way simpler that Qwen3.8 27B does not get; only GLM-5.2 and K3 do.
**
Analyze the 2 structural (not semantic) patterns in this text:
1. Morning revient.
2. Birds saluent Morgenlicht.
3. We suivons Waldwege toward maison.
4. Rain tombe plötzlich; we cherchons Schutz beneath sapins.
5. Night vient langsam; we trouvons Wärme near le Feuer, sharing quelques Geschichten together.
**
The answer should get not just the obvious cyclic E-F-G pattern but also the word counts being Fibonacci. Surprisingly few models get this. The only way I got Qwen to do this was on the 2.4T model, with extensive prompt scaffolding. Claude (Opus & Fable), Sol, Grok etc. got it on the first attempt. (All models, all attempts max reasoning level.)
This also seems like quite an esoteric use-case to me, but I guess some people might need to know when text follows a fibonacci sequence in terms of word count.
That problem sounds reminiscent of one I like to use as a benchmark, which is to request that the model create an .SVG of a logarithmic spiral of 50 numbered stones. Qwen 3.8 27B absolutely knocked that one out of the park, where a lot of larger models have failed outright or otherwise performed suboptimally.
Can you share an example of the Greek-history puzzle prompts you're talking about?
The stone remembers not what the cities gave, but what the goddess kept.
Begin with the first reckoning, under Ariston.
Find those who carried the Greeks in their name and the silver in their care. Take their name as we give it to them now, in capitals. Thirteen marks.
The goddess kept one from sixty. She asks the same of every mark: give each its ordinary alphabetic number, divide by sixty, and keep what she could not take.
Each remainder walks with the next. The last returns to the first.
The first of each pair turns upon itself; the second joins it; then comes the place where the pair began.
Eleven takes its fill. Keep what remains.
Raise eleven courses beneath the thirteen marks.
In course (r), beneath mark (i), add the course to what Eleven returned. Reduce again by Eleven. Cut the stone if the result is the remainder belonging to mark (i), or to the mark walking beside it.
Count courses from zero.
The mason turns where the stone turns: the first course goes with the writing, the next against it, and so on.
At the risk of sounding like a conspiracy theorist, this sounds like a great opportunity to make a statement. US or China, but likelier to be the former. Maybe Clem's on a call with the US government right now?
Open source teams have had access to the weights for at least a week now. vLLM folks expect full support on public release of the weights. Anything conspiratorial won't prevent the weights from leaking...
Well, sorry but look at the methodology on this. The study the post leans on did not study learning. 36 adults went around copying familiar words, and their "typing" is key presses using one index finger (!).
I'm afraid this one's close to going down the replication crisis whirlpool. According to Sol:
"Direct replications found no reliable learning advantage, while a 2024 meta-analysis found only a *small* handwriting benefit but a far larger typing advantage in information captured."
Typed Versus Handwritten Lecture Notes and College Student Achievement: A Meta-Analysis” by Abraham E. Flanigan and colleagues, published in Educational Psychology Review on July 12, 2024
The Greeks themselves, and us in the modern age, really have not given Egypt credit enough for it being the fount from where Greece drew to formulate much of its technology.
There is very little direct textual attestation of course, and I'm not advocating for some Black Athena type of research, but in this context of robotics, simply take note of these two artefacts:
Someone well informed on the literature might say, "Homer was dreaming of robots already in the 8th century BC with Hephaestus' workshop", and that might indeed seem like the earliest possible imagination of robotic creation. But note the sheer antiquity of the artefacts above, and if you apply a layer of very reasonable heuristics, it will become apparent that the Greeks were uniquely brilliant innovators, but they did so by brilliantly synthesizing and building on top of various neighboring civilizations and of course their own IE heritage.
Nope, GLM 5.2 is only the latest and greatest in a long line of open-weights models. There are even fully open source models that are comparable to o1-mini (OLMo), or almost-fully-open ones that are comparable to o3 (Nemotron).
I'm super grateful to the open labs (who, importantly, do not have the word 'Open' in their name), all the more so to the likes of Ai2.
There is no magic moat indeed. It is math, engineering and of course copious amounts of data (and the political maneuvering required to secure it, e.g. how most everyone has trained on Anna's Archive by this point).
Prompt: If A goes to B who then goes to C, can A send something to C?
Response:
We need to interpret best. The phrase "If A goes to B who then goes to C, can A send something to C?" could be a puzzle about the concept of sending something (like passing a ball) and the relationships.
Scenario: A gives something to B, and B passes it on to C.
Question: Can A also give the same thing to C?
Answer: Only if A can obtain a second copy (e.g., the thing was duplicated). Otherwise, after handing it to B, A no longer holds it and cannot “send” it unless a copy exists.
[Lots of other unnecessary commentary and "scenarios" that make even lesser sense]
Ran the same query and there is a ton of stuff, but it looks like it's reasoning through the ambiguity of the sentence. It still gets the right answer. Moreover, if we consider the FLOPs expended to get to the answer, and compare that to opus, I think it's still a net win.
My hunch is that Opus scale models probably have shortcuts encoded into the model that handle these ambiguities cases, wheres this model has learned a program to reason through the edge case (crystalized vs fluid intelligence). Remembering that probablity (frontier) vs calculating it on the fly (vibethink)
> LLM-based Query Quality Filtering. We utilize capable LLMs to assess query quality, filtering out samples with incomplete descriptions, unreasonable conditions, invalid logic, or an inability to effectively assess target knowledge points.
I have now witnessed this myself after not believing this at first. Of course, screenshots etc. will hardly prove anything. This needs a proper third-party audit!
reply