Hacker Newsnew | past | comments | ask | show | jobs | submit | cge's commentslogin

I try to do something like this with my publications, and encourage others to. My goal is to have the pipeline from raw data to complete figures and manuscript in a repository, with cached data for computationally expensive analysis and for stochastic simulation results, and the option for the user to just use those or run the full pipeline, with or without the same random seeds. I just make clear that the code was run-once code and is going to be messy compared to code refined over time and diverse uses. I generally use Zenodo to a GitHub repo, however, in case GitHub decides to do something bad in the future. Making sure things run far in the future can also be a challenge. Sure, you can use a container: will the base of that container be available in 30 years?

And with that said, for experimental work, this approach does not make things fully reproducible; it only makes the analysis reproducible. There are always factors that influence experiments: research is by definition at the edge of our understanding, and reality has countless variables, including ones no one has thought of, known about or thought important.


Caveat: I am not a researcher (yet), I am moving from coding into science via a new degree, and along the way I am helping troubleshoot bioinformatics pipelines for scientists.

I see a lot of reliance on containers to make code always available, and I have the same misgivings as you do. It'll work a few years into the future, but what happens once packages aren't compatible with each other/the base container is upgraded/etc.

I've already seen this with older bioinformatics code, which is on old repositories that aren't running anymore, or are very unreliable (but weren't at the time that the code was written). And I'm talking about code that's "only" 15 years old; people will be going to these papers for implementation details long after that point.

Zenodo seems like a good step in the right direction. It should be available as long as CERN is going, shouldn't it? And by that stage it should be "too big to fail".


You're goat, I'm trying to reproduce code from a paper and by following their instructions I can't even get packages to install because they conflict

When I had a Fairphone 4, regulars on the community forum repeatedly reminded everyone that repairability was not a primary goal of Fairphone, ethical consumerism was, and any longevity and repairability was secondary or coincidental.

Yes, that goes against most of Fairphone's own advertising, but it was consistent with my experience of the phone and discussions on the forum: security updates frequently so far out of date that some software could not be run, flimsy components inconsistent with advertising (yes, the battery was nominally swappable, but the snaps on the back would break easily, and support would claim that it was not intended to be removed regularly), parts that were frequently out of stock, and of course, whole new phone designs every generation rather than Framework-like component upgrades.

It seems like they may have improved since then, but those problems, along with the atrocious security (FP4 used AOSP's public test signing keys, with publicly-available private keys, for its firmware) and sketchiness around specs, standards and openness (especially for the camera), generally turned me off the company.


The entire paper is almost certainly Claude-generated, given the clear Claude-speak throughout. It seems like they didn’t even try to have Claude write a polished paper: it reads like Claude writing up a lengthy report, down to the needless sectioning with idiosyncratic title language. There appears to be no disclosure, and the author contribution section appears to falsely claim that a particular author wrote the text.

I know arXiv has taken some measures to combat spam like this, but it seems like they’ll need to do more. There’s just very little barrier now to creating giant slop papers like this and then dumping them anywhere that won’t reject them. It is an insult to everyone’s time, and I can’t imagine they expect people to actually read this. If the expectation is that everyone will use an LLM to interpret it, then maybe they should have at least had a few more rounds of tightening and polishing the paper, even via LLM, to save the redundant token use.


The only way I can interpret the percentage is that they are stating the increased cost as a percentage of sales tax rather than a percentage of the sale, such that "26% higher sales tax" in a state changing 10% sales tax would mean paying 2.4% more in total. That choice seems misleading, but does make the percentage make sense.


That is what it says after all, it's pretty explicit.


It's just such a bizarre choice that one might hope there would be another interpretation. Why measure a percentage change on sales tax, which varies heavily from location to location, and is not what the associated fees are based on, rather than simple choice of total cost?


Because sales tax is something you pay that's more than the sticker price, and people tend to have an intuition for sales tax in their area. Personally, I find "a 3% credit card fee is like paying 21% more in sales tax" to be intuitive.


That’s also very common in the UK, and shows up elsewhere in Europe sometimes. In the UK, they are legally but not practically or socially discretionary. There are some tax advantages for the restaurant and in the ideal case for staff.


> That’s also very common in the UK

Touristy places in central London maybe. They wouldn't dare try sneaking an undeclared service charge onto a locals bill in Northern England or Scotland :-)

> In the UK, they are legally but not practically or socially discretionary.

Speak for yourself. When entertaining clients at touristy places in London, one routinely asks for this underhand charge to be removed from the bill. (Whether or not we then choose to leave a cash tip is a matter of discretion, but the chances are rather low after such shenanigans).


IME they mostly only appear for large tables in the UK.


Then is should be declared in advance on the menu.

Adding an undeclared "service charge" to the bill only when it is presented to you is what we are talking about here. It is an increasingly common underhand tactic in and around London and an charge which one should ALWAYS ask to be removed ("that wasn't on the menu mate").


>In a world where things such as GLM-5.3, DeepSeek-V4-Pro-0813, and Kimi-K3 exists, this is a bit laughable.

This is not just for security, too. Using Fable for anything that could be remotely construed as being connect to chemistry or biology was impossible until a few weeks ago. Now it is slightly better, but still fails on many completely innocuous projects.

So as other models advance, Anthropic's sole frontier offering to entire academic fields remains an Opus that seems to get worse in capability each release. They're starting to become a joke in my field: at a conference a few weeks ago, one presenter laughed when I asked about his use of Fable and pointed out that it would downgrade if the letters 'd', 'n', and 'a' were anywhere near each other, which is not that far from my experience.


For some reason, despite these things being rather simple and concrete distinctions, all the reporting around the case keeps being confusing, vague, and contradictory. I had been rather convinced by the Ars Technica article [1] that Iron Mountain was just the data center operator and this was colocation of completely OSS-owned and managed equipment. Iron Mountain's statements there certainly seem to say that. There's the complexity here that, in general, Iron Mountain does apparently provide both data storage, and colocation.

[1]: https://arstechnica.com/information-technology/2026/08/pbs-s...


Part of the problem would be that we have "tech" journalists writing this who have never been on either side of the transaction directly as a colocation customer, or an ISP/datacenter/hosting company, and drafted/reviewed contracts for such services. Nor have they ever gone and like, personally laid hands on a 2U rackmount server in a cabinet in a colo.

I'm not sure that I could reasonably expect a "journalist" to have those qualifications, but they could at least attempt to interview a neutral third party in the colo industry who can explain the distinction.


They comment years after immediately closing the bug as wontfix, vaguely talk about users for expressing interest in features while not being willing to open a PR, when people in the thread had actually been asking, over a long period of time, whether they would accept a PR or were just opposed to the feature outright, and say they would welcome community contributions even though they think the feature is unnecessary. Then they immediately lock the issue and delete at least one comment.

I certainly interpreted the response as meaning that anyone actually trying to implement the feature would find the PR an exercise in frustration. It may be there were some cultural confusions about context and implication, but considering that they locked the issue, people weren't exactly encouraged to ask for clarification. And the developers have a sketchy enough reputation already from other incidents.


> It may be there were some cultural confusions [...] And the developers have a sketchy enough reputation already from other incidents.

Yeaah, both these points kind of makes it clear that both you and the author might have previous history with the history that goes beyond the messages that were referenced, and I don't have that context at all. I have no idea what "previous incidents" you might be referring to, so with that said I'll say that everything I've previously stated only been based on the text of that particular issue, nothing else. Based on the text from the issue alone, seems they're open to have the feature proposed as a PR to them, then if they typically reject any PRs, I have no idea about.


As an illustration of this, 100% of the detections in the trial in 2026 have been false positives. Scanning tens of thousands of faces for each deployment, they’ve had exactly one detection, which was false:

https://www.btp.police.uk/SysSiteAssets/media/images/british...


Yes, 1 is 100% of 1, but doesn’t that simply mean they need a bigger sample?


I think the point is that there is a distinction between calls that are intrusive spam or arguably scams but not already illegal, and calls where the content of the call is itself unambiguously illegal.

Bans on unsolicited calls, or do-not-call lists, can be effective for the former, as it moves the behavior from being legal to being illegal, while for the latter, they are already breaking laws, often in a more severe way, and technical measures or enforcement are more important.

A phone company trying to trick someone into paying more per month for something they don't need is likely to be worried about being fined if caught calling a number on a do-not-call list. A scammer pretending to be the tax office and demanding payment of supposedly unpaid taxes is going to be in trouble if caught regardless of rules about what numbers can be called.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: