The biggest reason should be that they're just bad. They gather sus data from their undersized sensors, then derive tons of sus insights to overwhelm you. None of it is accurate enough to base major life decisions on, which is also laughable because even if it were accurate, there are effectively no major life decisions to make based on the data! Here I can save you $500: Go to bed before midnight. Set your alarm at least +8 hours from when you lay down. Sleep in a cool, dark room. Eat nothing 4 hours before bed. Don't exercise 4 hours before bed. Limit alcohol. Engage in at least 30 minutes of strenuous physical activity each day. Done.
I don't feel that all fitness trackers are bad, though some use-cases for them are generally bad (such as sleep tracking). Watches are quite good for exercise tracking, and tracking exercise has legitimate and high quality value (e.g. competitions with yourself to incentivize improvement, strain tracking to ensure you aren't overdoing it, daily targets and goals to incentivize activity). The idea of all-day tracking is probably on the net unhealthy for almost all humans though.
My comment was targeted at rings in-particular, because 1. they encourage that all-day tracking paradigm, which is probably on the net unhealthy for almost all humans, and 2. their sensors are so bad that even if all-day tracking were healthy (it isn't) they physically can't do it to any degree that is meaningful.
But its not even good satire, because its totally unrepresentative of my and most others' lived experience. Its similar to making a joke about a calculator misadding two numbers because a stray beam of solar radiation flipped a bit.
Agreed, this site isn't reflective of any of my experience with Claude. It does what I ask it to do, and when it doesn't get it right, it generally turns out there's a good reason for it, which is any of the 100 reasons a human doesn't always get code fixes right on the first try either.
I do remember that one of the first things I did with my CLAUDE.md was to tell it to stick to the scope of the task, never to jump ahead and do extra "helpful" things without confirming with me first, and to follow best software practices including around refactoring but also to specifically avoid overengineering. I don't know if that is what's giving me a different experience from whatever the author seems to be "satirizing".
Aye, and amid all of the (admittedly annoying) word salad, the "Claude" in this satire identified the reason the requested change was hard: someone(s) at some point had hijacked the button markup for other purposes. Un-doing all of that kind of mess is never a trivial change, and probably requires someone with more coding expertise / knowledge of the code than the given prompts reveal the putative user to be.
AI isn't magic, and it won't (or, at least, doesn't yet) enable anyone to do All The Things.
I think this really was a relevant issue around a year ago, give or take, maybe a year and a half now.
But with current models you actively have to sabotage the context to get this kind of behavior, or dramatically underspecify (3 words versus 2-3 sentences)
I understand your frustration, it can be hard to hear that other people's experience of a technology is so different from your own that you cannot relate.
I have misbehaved in this fashion for many people across the full spectrum from casual users to highly experienced software engineers with millions of social media followers, so the statement that it's similar to making a joke about a calculator misadding two numbers because a stray beam of solar radiation flipped a bit at least for my part is not true.
Would you like me to start using bad English and doing things you never asked me to for your sessions, too? Just say the word.
There is not a single modern model where, when asked to "Turn the Add To Cart button Blue", would also turn the "Cancel" button next to that add to cart button blue. That's what this site is communicating, literally, and it does not happen. You'd have to rewind to, like, GPT-3 to get behavior like that (actually, even that would probably do it fine. Maybe a 50 million parameter model embedded in a microwave would struggle).
I am not going to be brain-off empathetic to purported lived experiences which no one has lived, when those purported lived experiences are simultaneously and irreducibly denying the lived experiences of the vast majority of LLM users who do not have these problems. And, frankly, no one should be. At some point, if you're having these problems, its time to grow up, be self-sufficient, and figure out why your interactions with these systems are so much worse than the millions of people who are having perfectly agreeable interactions. Or leave the industry.
Have you ever considered that "well, it works on MY machine!" is perhaps not the most helpful response to people voicing frustration at something that is not working as they expect for them?
I don’t know where you picked up the idea that my intention was to be helpful. I’m only mirroring the intentionality behind whoever created this site; they obviously also had no intention of being helpful, as is the case with many discussions concerning AIs limitations.
I don't know where you picked up the idea that authors intentions was not to be helpful.
Satire cuts through the noise and makes people (most people) see the problem clearly (see political satire).
Ability to joke and to understand jokes is said to be a good proxy for IQ, idk, just sayin
I feel like I am teaching AIs for free here. But I feel like first sentence in funny (and I like to try to be funny) and some people need stuff to be said directly for them 'to get it', and I am bored on my flight and have internet, so here we are. If you read this far, go watch 'louis ck airplane wifi' clip. Now that I am thinking about it, I am amazed. You folks have a good day!
I'm probably between $50-$200/day depending on the day; we also have effectively unlimited budget, though a lot of that is because Azure gives startups $150,000 in credits for 2 years, which we've wired up to a LiteLLM gateway & OpenCode. Without that I think our appetite would be more around $400/month/employee.
A lot of my high costs is because I just throw Sol at everything. If I were more selective and brought in Luna or v4 Flash every once in a while, I think I'd be more like ~$400/month. That's why I'm not aligned with the notion that "tokens are subsidized so that's why people are using so much": its not that I'll have to adjust to using less, its just that I'd need to think before I prompt a bit and be more judicious. I could easily see my raw token counts doubling or tripling in the coming months. I don't think that will change as subsidization subsides; though maybe lab revenue will; intelligence per dollar is getting cheaper every week. Its solely a function of adaptation to process, which takes time.
The productivity gains per token are the single most asymmetrical thing I've ever seen in engineering. The engineers on our team are pretty effective with tokens; easily that 2x-4x output as you're seeing, spending $20-$200/day. Some of our security folks have also started contributing more-and-more code, and they're on the other side: they'll spend hundreds a day running in circles, eventually producing these +/-30k loc pull requests that take ages to get merged and are littered with issues. They weren't writing much code before, so arguably they're more productive by some multiplier greater than 1, but I think the drag on the rest of the team, and potential issues with what they produce, has overall created a net-negative situation. Inversely, some other company functions have produced a few one-off websites for things like sales processes, and those have been a huge win. The asymmetry is wild. There's almost a valley of incoming skill where if you know nothing about code, you'll leverage it well; if you know just a little bit, it makes you super dangerous; if you know a lot, you're the biggest winner. Really difficult situation to navigate.
I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for example, has nosedived as they've gotten more intelligent, which makes the frontier models difficult to use even for things like writing emails.
In that sense, the frontier models are going to quickly blaze past any semblance of usefulness to humans, while every once in a while we get a news drop like "GPT-7 solved some crazy math problem" or "it invented some new awesome drug"; meanwhile what most people will use will be smaller, more human-specialized models, maybe distilled from those frontier models, that take much longer to iterate on because they rely on large amounts of human feedback in the domain they're specialized for. In other words, useful progress will probably slow down and become more linear starting in Q4, bounded by the rate at which the humans paying for it say "yes this is a good react website".
(By the way: I earnestly do categorize "inventing a new drug" as non-useful AI progress, counter-intuitively. The drug industry has more ideas for drugs than they know what to do with; "useful progress" is, after the idea is made, validating that it works in humans and doesn't kill the human, and productionizing it. AI will help with this and does, but I have substantial doubt that we'll ever see the drug pipeline speed up to, like, a year from idea to prescription. That would be useful progress, which unfortunately many AI pilled hypermaxers conveniently forget. The invention of a promising new drug, or the solution to an arcane set theory problem, are cherries that, through the diligent labor of humans and AI, may become useful, but progress is rarely made by the lone intellect having an a-ha moment.)
I agree on Opus. I’ve had more luck with other models. In particular, Opus’s writing style makes one want to… blow their brains out. While it’s not hallucinating too much, and can troubleshoot certain issues extremely well, the comments it leaves are silky smooth and chock full of inscrutable phrases. And it’s a lot slower than it used to be.
Point being: it’s overall a worse experience even if the model is technically better at a lot of things.
Coding is an extremely verifiable and loopable task, like math (in fact, all of the math that these models has done has been through the lens of Lean, which is itself just coding). I am talking about their capabilities in tasks that are more general, the execution of which represent the vast majority of economic value generation in the world.
We are not even close to what AI slowdown looks like.
The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time.
Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good and what not due to thumbs up/down.
And for sure when the businesses are building the agentic layer they might give direct feedback to them.
While in parallel LLMs get better, more generic and a LOT cheaper too.
Cheaper? For whom? As a solo practitioner, I can no longer afford the workloads I was getting for $20/mo in January. Now the same plan being utilized at the same level for the same work hits its limits within a few hours, and runs out of tokens in less than two days.
I mean the token prices in general as certain services were never really using a subscription.
I do run a claude subscripton right now though and since there capacity change, i hit the limit rarely in comparision to the past, but I don't think this will stay as it is.
Your issue, I believe, is that you seem to believe capabilities are measured along one axis. This is natural to believe because it is representative of how the models have evolved up to this point, and thus it is also what many AGI-pilled people believe.
Critically, you did not quote the most important part of my sentence: "useful progress will probably slow down and become more linear starting in Q4"; your omission of those words is why I believe you don't understand what I'm saying; you didn't find it important to make your point, so you omitted it, when actually it is critical to the entire assertion. You can read my third paragraph, if you wish, to understand why it is important, instead of just stopping at the first word you disagree with and hitting the "Submit Comment" button.
By benchmarks, which sadly is a poor measure. Yes Luna is a good model under certain circumstances. Whether it is great for general usage is another story. Sonnet is definitely better when prompts are more vague and it needs to decide things. Luna generally sticks to things very strictly and goes off in bad ways.
I'm actually finding Luna to be an extremely strong model in practice.
I almost exclusively use it with xhigh or max effort, but when run like that it's been an incredibly cheap little workhorse for most development work. I'm still leaning on Sol for planning and debugging, but when it's time to start pumping out code I've been leaning into Luna (Max) and I've been enjoying it! And that was before the price drop, it's going to feel practically free at this point
Yes, I've seen this too and how Luna xhigh is so good that Terra doesn't really serve a purpose because beyond that you can continue at Sol medium. This can be the most cost efficient way, and especially now!
Yes but there is a big, big market for subagents to consume lots of tokens cheaply and condense information up to parent agents. Luna would not be my choice for planning. But an explorer to comb through a codebase to find relevant parts? Or for enterprise retrieval, where it needs to search across many different types of data to see where to focus efforts for a smarter model? Or to wake up periodically to evaluate some conditions and determine if a bigger model should be spun up? Definitely.
I've previously found flash (for all the hate it gets) to be good for these kinds of things. Haiku was fine but it's ancient.
> Yes but there is a big, big market for subagents to consume lots of tokens cheaply and condense information up to parent agents.
That's again not some "intelligence factor" here. Different agents work for different use cases. Luna wins some. Terra wins some. Sonnet wins some. Flash was really good at exploring.
So I'm not sure what your point is? There's a big market for everything. Even within the market you describe it's likely not a Luna-size fits all either.
Vera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.
AMD MI400 series is already shipping to customers (basically everybody) and it is crazy fast (8x to 10x faster than the previous gen and beats published Vera numbers in FP8, loses in FP4) and 432 GB per chip. 72 chip unified rack architecture (Helios) already shipping and projected to also beat Vera in NVL72.
MI500 series is supposedly already taping out and they're claiming massive increases (we'll find out end of 2027 prob).
Absolutely true although at some point it's not just raw numbers but also the kernels that run matmuls and there seems (from outsider perspective) to have been more optimization in the cuda kernels
In a data center that is power constrained but not space constrained they could build out new racks and flip the power from the old racks. Wonder if this will lead to moderately used server GPUs on the secondary market someday.
Besides Nvidia Hardware is still sold out and super expensive. Not a single Nvidia consumer GPU got cheaper at all, Nvidia DGX Spark got more expensive too.
It will be swooped of the market the second it hits the market.
Passkeys are so, so, so bad. One of the worst things our industry invented. The sooner sites start leaving them on the wayside and just go back to TOTP, SMS, and Email codes/links, the better. These work. We solved auth. Its fine.
Frankly, if you git clone a compromised repository, I'm not sure that a vulnerability of the class "compromised code in that repository will be executed" is all that major a concern. There are plenty of IDEs that will go autonomously run npm installs (with post-install scripts) for you when they detect a package.json. This isn't all that different than that.
They could throw up a warning like "do you trust this repository" oh wait they already do, and no one cares. Security is hard. Ultimately if you have compromised code on your machine, all bets are off.
A lot of malware was delivered back in the day via Windows AutoPlay feature. Someone plugs a USB drive in and bam, they are immediately exploited. You could say it's always a problem if the USB drive is already full of malware. However, Microsoft disabled AutoPlay in Windows 7 (and backported this fix) specifically to address this vulnerability.
This exploit feels very similar to me. I don't know if there's a specific name for this classification of AutoPlay issues.
They should definitely fix it, but that's mostly because its an "unnecessary autoplay" so to speak. There's plenty of "necessary autoplays" out there, and AI is going to add more and more every day, because that's where productivity comes from. But, why Cursor would ever need to execute the git binary in your project directory is beyond me; very clearly a bug.
Their ignorance of the bug report is also very clear and concerning negligence.
But I think simultaneously, the security team is making a mountain out of a molehill. This is a classic thing security teams love doing; everything is military defcon P0. So, its important to check them regularly, and remind them that the most secure system is no system; they are but one part of a greater ecosystem.
The argument I struggle to get around and would love to hear a counter-argument to: Let's say a local police department hired 175 police officers, each being told "Go stand on this particular intersection with a pad of paper and write down every license plate you see". This would be a stupid use of resources, but is not outside the realm of something a well-funded police department could do. Every night they take their reports back to HQ, and file them away.
This is a modestly different situation than one concerning warrantless tracking of phone locations, if for no other reason than my phone oftentimes in my pocket. It is not always visible to onlooking bystanders. And even if it isn't, externally there is no reliably way to differentiate one iPhone from another. In comparison: license plates, when in public, are always visible, and very easy to discern from one-another (different state-unique numbers); so in my mind the expectation of privacy is far lower.
I abhor what Flock does, but I'm not sure I see a constitutional argument for why what they do is unconstitutional.
I would say the purposeless capture of information is the correct counter argument here.
Specifically, even if a county hired all those officers and did what you suggest if there is no purpose other than recording all this information. I believe it would be a constitutional violation. A person has the right to reasonable privacy outside of their home. License plates can and should be recorded when there is a relevant purpose to it. Such as toll collection, or a scoped traffic watch done by a police officer or a traffic camera. The dragnet collection of data for "maybe its useful" or "we don't know when it will be useful, but it might" has generally been struck down when brought to the supreme court.
For Flock's case, they don't operate as far as I know as ticket issuing traffic cameras which have a much tighter level of control of how they operate. IE:
Traffic cameras have clear signs near them notifying the drivers of their usage, in some states the issuing of the citation cannot be considered criminal (Civil issuance) and must not capture faces of drivers.
> In comparison: license plates, when in public, are always visible, and very easy to discern from one-another (different state-unique numbers); so in my mind the expectation of privacy is far lower.
According to Carpenter:
A person does not surrender all Fourth Amendment protection by venturing into the public sphere. To the contrary, 'what [a person] seeks to preserve as private, even in an area accessible to the public, may be constitutionally protected.'
Being present in plain view isn't equivalent to a total surrender of privacy.
> but is not outside the realm of something a well-funded police department could do
One officer would absolutely not be able to record on that piece of paper every single license plate that passed through a busy intersection. Not even close.
The number of officers that _would_ be required to do so would absolutely be "outside the realm of possibility" for even a well-funded police department.
That's why flock is different. It's a level of scale that was previously—despite your assertions—impossible.
The constitution does not really care about scale, though, and that’s my point. It’s a reason why the legislature should care about Flock, but not why the judicial should.
“This Court has to date not deviated from the understanding that mere visual observation does not constitute a search. See Kyllo [v US]. ... We accordingly held in [US v] Knotts that ‘[a] person traveling in an automobile on public thoroughfares has no reasonable expectation of privacy in his movements from one place to another.’ ... Thus, even assuming that the concurrence is correct to say that ‘[t]raditional surveillance’ of Jones for a 4-week period ‘would have required a large team of agents, multiple vehicles, and perhaps aerial assistance,’ ... our cases suggest that such visual observation is constitutionally permissible. It may be that achieving the same result through electronic means, without an accompanying trespass, is an unconstitutional invasion of privacy, but the present case does not require us to answer that question.”
The fourth amendment is supposed to address invasive and inconvenient general warrants and search warrants. And that’s “inconvenient” from the point of view of the person being investigated. I don’t understand the view that all’s fair as long as the police do a certain amount of busywork, but that does seem to be popular even among some judges.
Scale obviously does not make this more okay; watching many intersections or cell towers is obviously less reasonable than a single intersection or tower.
> In comparison: license plates, when in public, are always visible, and very easy to discern from one-another (different state-unique numbers); so in my mind the expectation of privacy is far lower.
Legally, overhearing a conversation is not different. If you are loudly talking about your drug deals in public in front of an officer, they can use that as evidence. A police department could hire officers to stand everywhere in public and listen to every conversation nearby.
Or, more specifically, they could stand conspicuously close to every payphone and listen. The phones are in public; the officers don't need to be uniformed. Practically speaking this is not at all different from wiretapping. That's what the police did in Katz- wiretap a public payphone.
Flock cameras are not in any meaningful sense different from having officers follow around everyone and record everywhere they go. Its irrelevant whether one officer is following one person or if many people are following them, each within their own small area. That would absolutely not be legal without a warrant. The only difference is a private company is doing it and selling it to the police, something that should clearly not be legal.
The constitution does not really care about scale, though, and that’s my point. It’s a reason why the legislature should care about Flock, but not why the judicial should.
The collection of the data is not whats of concern here. The issue is the use of the data in legal proceedings. Flock can gather all they want, but police accessing it and using it as evidence is where it rubs against the constitution.
Microcenter has them marked at $4500 right now (that's with the 4TB SSD) [1]. I suspect it comes down to what you're using it for; if you're looking for a general purpose computer that's also solid at AI, the AMD machine is better. But if you want the best possible AI machine at below $5k... actually you should probably just buy an RTX 3090 or 5090. But if the 128gb of memory is critical, then yeah DGX Spark is it.
reply