Top Stories
GLM 5.2 beats Claude in our cyber benchmarks
879 points · semgrep.dev
Semgrep ran an open-source Chinese model, GLM 5.2, against Claude on their internal cybersecurity benchmarks and found it came out ahead — a result that lands hard given how much the frontier-model conversation has assumed a US lab lead. The post is careful about methodology, which is why HN is engaging rather than dismissing it: the interesting question isn’t just the score, but how quickly open-weight models are closing the gap on specialized, high-value tasks like vulnerability reasoning.
Age verification is just a precursor to automated attribution of speech
546 points · nonogra.ph
A sharp argument that mandatory age-verification regimes aren’t really about protecting kids — they’re the infrastructure for tying every online utterance back to a real identity. The piece resonates with HN’s long-standing privacy instincts, framing ID checks as the thin end of a wedge toward end-to-end attribution of speech. It pairs naturally with today’s KIDS Act story below.
HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74. No – 88
510 points · danunparsed.com
HackerRank open-sourced its applicant-tracking system, and the author promptly fed his own resume through it — getting wildly different scores on repeated runs. It’s a vivid, funny demonstration of how non-deterministic and arbitrary automated resume screening can be, and HN is seizing on it as evidence that the gatekeepers rejecting candidates may be little more than a coin flip dressed up as a score.
The KIDS Act would require age checks to get online
504 points · eff.org
The EFF breaks down proposed legislation that would mandate age verification across much of the web, warning about the chilling effects and surveillance costs that come with checking IDs at the door of the internet. Coming from the EFF, this is catnip for HN — and read alongside the attribution-of-speech essay above, it captures a growing unease that “for the children” policy is quietly building an identity layer over the open web.
I used Claude Code to get a second opinion on my MRI
447 points · antoine.fi
A developer walks through feeding their MRI imaging to Claude Code (Opus) for a layperson’s second read, and the result is equal parts impressive and unnerving. HN’s debate is the obvious one: how much trust should anyone put in an LLM interpreting medical scans, where the model’s confidence and its accuracy may diverge in exactly the cases that matter most. It’s a great snapshot of how people are actually reaching for these tools in high-stakes personal moments.
Professor denounces mass AI fraud on an exam at Brown
421 points · english.elpais.com
A Brown professor goes public about widespread AI-assisted cheating on an exam, reigniting the now-perennial argument about what assessment even means in the LLM era. HN commenters split predictably between “the exam format is broken, not the students” and “this is straightforward academic dishonesty” — but everyone agrees the old equilibrium is gone and nobody has a clean replacement yet.
Librepods: AirPods liberated
404 points · github.com
An open-source project that reverse-engineers AirPods to bring features like ear-detection, battery status, and spatial controls to Android and Linux. This is classic HN catnip — interoperability won back from a walled garden through clean reverse engineering. Beyond the cool factor, it’s a small data point in the ongoing fight over whether the hardware you bought should only work the way the vendor intends.
5k menus from the New York Public Library’s Buttolph Collection (1880-1920)
375 points · pudding.cool
The Pudding turns a digitized archive of roughly 5,000 historical restaurant menus into a gorgeous data-driven story about how Americans ate, drank, and priced their meals across four decades. HN loves this genre — primary-source data, thoughtful visualization, and a window into everyday economic history — and it’s a nice palate cleanser amid the AI-policy churn.
Historical memory prices 1960-2026
321 points · stanford.edu
A Stanford-hosted dataset charting the cost of computer memory across more than six decades, a staggering decline that underpins essentially everything in modern computing. The HN discussion digs into the inflection points, the plateaus, and what the curve implies for an era where ML workloads are once again making memory the bottleneck.
Also Trending
- A way to exclude sensitive files issue still open for OpenAI Codex (210 points) — A long-open GitHub issue asking Codex to respect a way to exclude sensitive files keeps surfacing security concerns about agentic coding tools. github.com
- Boeing 747 begins its final descent (194 points) — The Atlantic on the slow retirement of the jumbo jet that reshaped global travel. theatlantic.com
- Pollen’s CEO tried to remove a critical article, and Google helped (183 points) — Gergely Orosz on a startup attempting to scrub unflattering coverage via takedown requests. blog.pragmaticengineer.com
- Model Training as Code (165 points) — Aleph Alpha makes the case for treating model training pipelines as declarative, version-controlled code. aleph-alpha.com
- Tokenmaxxing is dead, long live tokenmaxxing (157 points) — A wry look at how the economics of squeezing tokens keeps reinventing itself in the agentic era. 12gramsofcarbon.com
- TOP500 at ISC’26: We have a new Number 1 supercomputer (111 points) — Chips and Cheese covers a fresh leader atop the supercomputing rankings. chipsandcheese.com