Skip to main content

Share story

Science AI

New Benchmark Tests Whether LLMs Can Reconstruct Expert Investor Decision Frameworks

A new benchmark called InvestPhilBench evaluates whether large language models can accurately reconstruct and apply the procedural reasoning of expert investors. The v0.6 release contains 118 verified principle cards, 25 decision-framework cards with topology metadata, and 243 questions split between development and held-out test sets. An automated scoring pipeline called BASP and a gate-level metric called Gate Reconstruction Accuracy measure whether models follow the correct reasoning steps rather than just reaching the right answer. A preliminary four-model evaluation on the development set showed a sharp tier split: frontier models scored 0.906 on the composite metric while mid-tier models scored 0.438. However, even the best model achieved only about 0.77 on gate-level reconstruction for mid-level reasoning and 0.57 to 0.62 for advanced extrapolation tasks, suggesting fluent output can mask procedural gaps.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from arXiv and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Science
AI Science

Preprint reports schema-free generation of valid enterprise test data

Researchers report in an arXiv preprint that their Generalist Populator agent generated enterprise data with 100% constraint satisfaction and 0.88 average marginal fidelity across ten simulated environments without accessing datab...

AI Science
AI Science

Preprint reports task-completion gains from agent-generated interfaces

Researchers report in an arXiv preprint that training an agent to generate interactive interfaces improved a 4B model's Pass@3 task-completion score from 9.33% to 58.00%. Their GenUI-Harness pairs an agent that retrieves informati...

Robotics AI
Robotics AI

DreamTrue researchers report fewer interaction defects in robot video predictions

Researchers report that DreamTrue, a model that predicts videos of robot actions, reduced human-assessed interaction defects from 48.12% to 6.25% on AgiBot. The preprint addresses predictions that follow commands inaccurately or f...

Robotics AI
Robotics AI

FAITH preprint reports humanoid safety gains while preserving task performance

Researchers report that their FAITH safety filter achieved a 99.95% safety rate while retaining 97% of unfiltered task return on a 29-degree-of-freedom humanoid in Walking-Avoid. The preprint also describes demonstrations of the s...

AI Science
AI Science

AMD, OpenAI and Google join AI resource pledges for federal science

AMD, OpenAI and Google are pledging AI tools and compute credits for the federal Genesis Mission alongside Nvidia and Anthropic, The Next Web reported. AMD's commitment is $500M, OpenAI's is $200M and Google's is $150M. The pledg...