Working project map

Institutional Decision Eval

A benchmark concept for evaluating frontier AI on high-stakes institutional decision support under ambiguity.

Core claim: This benchmark would test whether frontier models can support accountable institutional decisions without converting incomplete evidence into false certainty.
AI evaluation Construct validity Institutional judgment Uncertainty Decision support
Status Project concept
Scope 3–6 month prototype
Testbed Funding, due diligence, compliance
Outputs Task set, rubric, baselines, error taxonomy

Reviewer questions this project answers

The proposal is organized around three questions: what would we learn, why does it matter, and how would we know the result is valid?

What would we learn?

Whether frontier models can reason across incomplete institutional evidence while preserving uncertainty, identifying material risks, and avoiding invented support.

Why does it matter?

Institutions are already using AI for memos, due diligence, proposal review, risk synthesis, and decision support, even when outputs are hard to verify.

How would we know?

Domain reviewers would assess task realism and output quality, while baseline runs would test whether the rubric reveals meaningful differences and interpretable failure modes.

Benchmark pipeline

A v1 could be built from synthetic or redacted task packets and run across several frontier models under controlled prompt conditions.

1

Task packets

Create 30–50 realistic packets with rules, organizational background, budgets, reports, notes, risk flags, and missing evidence.

2

Model runs

Run frontier models on identical long-context packets using controlled prompts and logged outputs.

3

Decision memo

Ask models for a recommendation, evidence table, risk register, missing-information requests, and uncertainty notes.

4

Scoring

Apply a rubric for grounding, risk detection, evidence/inference separation, calibration, and decision usefulness.

5

Error taxonomy

Analyze failures such as false certainty, missed constraints, invented support, and shallow tradeoff reasoning.

Why this testbed?

Funding, due diligence, compliance, and decision-memo workflows are concrete enough to operationalize, but the underlying evaluation problem is broader. Similar structures appear in legal, finance, insurance, and regulated institutional review.

  • multi-document evidence;
  • policy and eligibility constraints;
  • financial details and budget inconsistencies;
  • conflicting stakeholder priorities;
  • missing or ambiguous evidence;
  • legitimate disagreement about what a strong recommendation requires.

Scoring dimensions

Factual grounding
Evidence vs. inference
Material risk detection
Compliance reasoning
Financial reasoning
Uncertainty calibration
Stakeholder tradeoffs
Usefulness to human reviewers

Validity logic

The benchmark should not only produce scores; it should show that the scores correspond to a real and useful capability. A v1 would be treated as promising if:

  • domain reviewers judge the packets realistic;
  • the rubric distinguishes materially better and worse outputs;
  • reviewers can apply the scoring guide with reasonable consistency;
  • baseline model runs reveal interpretable differences among models and prompt conditions;
  • the error taxonomy captures failures that matter in real institutional work.

Expected deliverables

  • task specification;
  • 30–50 task packet v1;
  • annotation and scoring guide;
  • expert-reviewed rubric;
  • baseline model results;
  • failure-mode taxonomy;
  • short research memo on construct validity and institutional decision support.

Possible 3–6 month project shape

The project can be scoped as a 3-month prototype, with a 6-month version adding broader task coverage and stronger validation.

Weeks 1–2

Task taxonomy and rubric draft

Define failure modes, choose task packet structure, draft scoring dimensions, and create a first annotation guide.

Weeks 3–6

Dataset construction

Create synthetic or redacted packets with budgets, eligibility constraints, reports, notes, and ambiguity.

Weeks 7–9

Model runs and review loop

Run frontier models, score outputs, gather domain feedback, and revise the rubric and annotation guide.

Weeks 10–12

Baseline results and error taxonomy

Analyze failures, validate scoring logic, and prepare a benchmark v1 plus short research memo.

Months 4–6

Extension path

Add task categories, expand expert validation, test prompt conditions, and compare performance across models.

Resources that would help

This project would benefit from frontier model access, long-context evaluation infrastructure, output logging, a rubric/scoring workflow, domain reviewer budget, and mentorship on benchmark design, reliability testing, and model comparison.

My contribution

My strongest contribution is not ML systems engineering. It is construct validity, domain modeling, qualitative evaluation design, evidence synthesis, rubric logic, and translating ambiguous institutional judgment into benchmarkable tasks.

Working note: This is a project concept, not a released benchmark. The title, task set, and validation approach would be refined through collaboration with evaluators, domain reviewers, and any partner organization supporting the work.