What would we learn?
Whether frontier models can reason across incomplete institutional evidence while preserving uncertainty, identifying material risks, and avoiding invented support.
A benchmark concept for evaluating frontier AI on high-stakes institutional decision support under ambiguity.
The proposal is organized around three questions: what would we learn, why does it matter, and how would we know the result is valid?
Whether frontier models can reason across incomplete institutional evidence while preserving uncertainty, identifying material risks, and avoiding invented support.
Institutions are already using AI for memos, due diligence, proposal review, risk synthesis, and decision support, even when outputs are hard to verify.
Domain reviewers would assess task realism and output quality, while baseline runs would test whether the rubric reveals meaningful differences and interpretable failure modes.
A v1 could be built from synthetic or redacted task packets and run across several frontier models under controlled prompt conditions.
Create 30–50 realistic packets with rules, organizational background, budgets, reports, notes, risk flags, and missing evidence.
Run frontier models on identical long-context packets using controlled prompts and logged outputs.
Ask models for a recommendation, evidence table, risk register, missing-information requests, and uncertainty notes.
Apply a rubric for grounding, risk detection, evidence/inference separation, calibration, and decision usefulness.
Analyze failures such as false certainty, missed constraints, invented support, and shallow tradeoff reasoning.
Funding, due diligence, compliance, and decision-memo workflows are concrete enough to operationalize, but the underlying evaluation problem is broader. Similar structures appear in legal, finance, insurance, and regulated institutional review.
The benchmark should not only produce scores; it should show that the scores correspond to a real and useful capability. A v1 would be treated as promising if:
The project can be scoped as a 3-month prototype, with a 6-month version adding broader task coverage and stronger validation.
Define failure modes, choose task packet structure, draft scoring dimensions, and create a first annotation guide.
Create synthetic or redacted packets with budgets, eligibility constraints, reports, notes, and ambiguity.
Run frontier models, score outputs, gather domain feedback, and revise the rubric and annotation guide.
Analyze failures, validate scoring logic, and prepare a benchmark v1 plus short research memo.
Add task categories, expand expert validation, test prompt conditions, and compare performance across models.
This project would benefit from frontier model access, long-context evaluation infrastructure, output logging, a rubric/scoring workflow, domain reviewer budget, and mentorship on benchmark design, reliability testing, and model comparison.
My strongest contribution is not ML systems engineering. It is construct validity, domain modeling, qualitative evaluation design, evidence synthesis, rubric logic, and translating ambiguous institutional judgment into benchmarkable tasks.