Kurate preprint demonstrates evidence-quality scoring across 4,347 papers
Researchers report in an arXiv preprint that Kurate, a system using large language models, assessed study quality across 4,347 papers, including 3,913 reporting randomized trials. It scores eight dimensions of study design and rep...
MIT flight planner avoids moving obstacles in 12 UAV tests
MIT researchers developed a flight planner that avoided all dynamic obstacles in 12 real UAV test flights, using onboard computing and sensors to replan trajectories. Called SANDO, the system also reached its goal faster than seve...
SchemaFill preprint reports up to 4.05× throughput for AI tool calls
Researchers report that SchemaFill improved end-to-end throughput by up to 4.05× over token-by-token generation on the Glaive and BFCL benchmarks. The preprint presents a method for generating structured calls that let language mo...
Scheduling planner improves integration of parallel coding agents in preprint
Researchers report that a planner built into Nerveplane increased clean integration from 1/9 to 9/9 scenarios and reduced merge conflicts from 13 to 0 in a benchmark of parallel coding agents. The preprint evaluates coordination t...
OpenAI claims internal model resolved more than 100 open math problems
OpenAI says an internal model has resolved more than 100 longstanding open problems across most areas of mathematics, alongside the Navier-Stokes Millennium Prize problem. Spokesperson Lindsay McCallum told WIRED that training beg...
OpenAI publishes 722 math manuscripts from an unreleased model
OpenAI said it released 722 mathematical manuscripts organized into 372 families in a public GitHub repository on October 6. The release adds proof artifacts and 10 abridged reasoning summaries to results produced by an unreleased...