Top Stories

GLM 5.2 beats Claude in our cyber benchmarks

879 points · semgrep.dev

Semgrep ran an open-source Chinese model, GLM 5.2, against Claude on their internal cybersecurity benchmarks and found it came out ahead — a result that lands hard given how much the frontier-model conversation has assumed a US lab lead. The post is careful about methodology, which is why HN is engaging rather than dismissing it: the interesting question isn’t just the score, but how quickly open-weight models are closing the gap on specialized, high-value tasks like vulnerability reasoning.


Age verification is just a precursor to automated attribution of speech

546 points · nonogra.ph

A sharp argument that mandatory age-verification regimes aren’t really about protecting kids — they’re the infrastructure for tying every online utterance back to a real identity. The piece resonates with HN’s long-standing privacy instincts, framing ID checks as the thin end of a wedge toward end-to-end attribution of speech. It pairs naturally with today’s KIDS Act story below.


HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74. No – 88

510 points · danunparsed.com

HackerRank open-sourced its applicant-tracking system, and the author promptly fed his own resume through it — getting wildly different scores on repeated runs. It’s a vivid, funny demonstration of how non-deterministic and arbitrary automated resume screening can be, and HN is seizing on it as evidence that the gatekeepers rejecting candidates may be little more than a coin flip dressed up as a score.


The KIDS Act would require age checks to get online

504 points · eff.org

The EFF breaks down proposed legislation that would mandate age verification across much of the web, warning about the chilling effects and surveillance costs that come with checking IDs at the door of the internet. Coming from the EFF, this is catnip for HN — and read alongside the attribution-of-speech essay above, it captures a growing unease that “for the children” policy is quietly building an identity layer over the open web.


I used Claude Code to get a second opinion on my MRI

447 points · antoine.fi

A developer walks through feeding their MRI imaging to Claude Code (Opus) for a layperson’s second read, and the result is equal parts impressive and unnerving. HN’s debate is the obvious one: how much trust should anyone put in an LLM interpreting medical scans, where the model’s confidence and its accuracy may diverge in exactly the cases that matter most. It’s a great snapshot of how people are actually reaching for these tools in high-stakes personal moments.


Professor denounces mass AI fraud on an exam at Brown

421 points · english.elpais.com

A Brown professor goes public about widespread AI-assisted cheating on an exam, reigniting the now-perennial argument about what assessment even means in the LLM era. HN commenters split predictably between “the exam format is broken, not the students” and “this is straightforward academic dishonesty” — but everyone agrees the old equilibrium is gone and nobody has a clean replacement yet.


Librepods: AirPods liberated

404 points · github.com

An open-source project that reverse-engineers AirPods to bring features like ear-detection, battery status, and spatial controls to Android and Linux. This is classic HN catnip — interoperability won back from a walled garden through clean reverse engineering. Beyond the cool factor, it’s a small data point in the ongoing fight over whether the hardware you bought should only work the way the vendor intends.


5k menus from the New York Public Library’s Buttolph Collection (1880-1920)

375 points · pudding.cool

The Pudding turns a digitized archive of roughly 5,000 historical restaurant menus into a gorgeous data-driven story about how Americans ate, drank, and priced their meals across four decades. HN loves this genre — primary-source data, thoughtful visualization, and a window into everyday economic history — and it’s a nice palate cleanser amid the AI-policy churn.


Historical memory prices 1960-2026

321 points · stanford.edu

A Stanford-hosted dataset charting the cost of computer memory across more than six decades, a staggering decline that underpins essentially everything in modern computing. The HN discussion digs into the inflection points, the plateaus, and what the curve implies for an era where ML workloads are once again making memory the bottleneck.