# New Benchmark Tests Whether LLMs Can Reconstruct Expert Investor Decision Frameworks

_Friday, July 10, 2026 at 12:06 PM EDT · Science, AI · Latest · Tier 1 — Major_

A new benchmark called InvestPhilBench evaluates whether large language models can accurately reconstruct and apply the procedural reasoning of expert investors.

The v0.6 release contains 118 verified principle cards, 25 decision-framework cards with topology metadata, and 243 questions split between development and held-out test sets. An automated scoring pipeline called BASP and a gate-level metric called Gate Reconstruction Accuracy measure whether models follow the correct reasoning steps rather than just reaching the right answer.

A preliminary four-model evaluation on the development set showed a sharp tier split: frontier models scored 0.906 on the composite metric while mid-tier models scored 0.438. However, even the best model achieved only about 0.77 on gate-level reconstruction for mid-level reasoning and 0.57 to 0.62 for advanced extrapolation tasks, suggesting fluent output can mask procedural gaps.

## Sources

- [arXiv](https://arxiv.org/abs/2606.25984)

---
Canonical: https://techandbusiness.org/newswire/A9KxU337ELsycETkom6no5
Retrieved: 2026-08-24T20:37:41.991Z
Publisher: Tech & Business (techandbusiness.org)
