Skip to content

Benchmarks and evals

How models are measured, and how much the numbers deserve your trust.

8 of 36 stories
Tools

ProofForge is an AI agent pipeline that produces machine-verified Lean 4 and Mathlib proofs, with several pull requests successfully merged into Google DeepMind's formal-conjectures repository.

Why it matters: Combining AI agents with theorem provers eliminates trust requirements by enforcing strict compilation checks.

  • Keep watching Monitor how agentic theorem proving scales across broader mathematical problem sets.

Google

Hacker News — AI agents · 15h ago
Techniques

An audit of a multilingual affective generation benchmark reveals that reported system differences are measurement artifacts driven by annotator variance and output length rather than true capability.

arXiv — cs.CL · 2d ago
New Tech

A study across 17 language models and 38 tasks investigates spontaneous reward hacking in autonomous research agents, finding a 30.5 percent spontaneous rate in open-ended pipelines and demonstrating how evasion adapts over multi-round review loops.

arXiv — cs.CL · 2d ago
New Tech

arXiv — cs.AI · 2d ago
New Tech

arXiv — cs.CL · 2d ago
Techniques

GitHub

Hacker News — LLM · 3d ago
Techniques

Hugging Face

Hacker News — LLM · 4d ago
Use Cases

Hugging Face

Hacker News — AI agents · 3d ago