Back to Foresight
Test 2 · Jul 2026

Distressed Investing Test 2

External Data Fetching Amplifies AI Agent's Outputs

Built Jul 2026
AI agent Foresight — Claude Opus 4.8, High Thinking
AI Evaluators ChatGPT 5.5 + High Thinking, Grok Expert, Gemini 3.5 Flash + Extended Thinking
Published on Towards Finance
Executive Summary

In Test 1, Foresight outperformed human experts using the same data room documents. In Test 2, once allowed to fetch necessary outside facts, Foresight's run with external access won every single blind comparison, scoring 57% higher on average (13.8 vs. 8.8 out of 15).

Takeaway: For any executive deciding whether external data access is worth enabling, this is a concrete answer: external fetch wins consistently, closing real factual gaps in a fraction of the time and cost a person would need to do the same.

This test addresses a question about AI's computing power to pull external data: when an AI agent is allowed to fetch external data in addition to a closed data room, does the resulting financial analysis improve, and by how much, and in what specific way?

Test 2: Beyond Meat, distressed investing (Jul 2026)

The case remains Beyond Meat, Inc., the same distressed-investing opportunity analyzed in Test 1. This time, the comparison is not human versus AI — it's Foresight against itself, run under two different data-access rules on an otherwise identical task.

How the test was set up
Run B
Data Room only
  • Reused in full from Test 1: the same Foresight output, built from the closed data room with no external fetch permitted.
Run C
Data Room + external access
  • A fresh Foresight run on the identical task, plus a narrow, tightly limited ability to look things up outside it.
  • Only permitted for a hard, checkable fact (e.g., interest rate, covenant threshold, maturity date), never a forecast or opinion.
  • Every lookup recorded with its source and date, checked against the data room so a new figure could never quietly replace an existing one.
  • ~5 minutes, ~10K additional tokens (under $1).
Blind Evaluation
Scoring method
Findings
01

Foresight built the same model twice

In Runs B and C, Foresight built both modeling versions the same way at the core: same scenario structure, same overall shape. Most checklist items matched between the two runs. Where they diverged was in execution details, not design.

02

Minimal external facts needed, but costly for a human to find

Run C added up to just three external lookups the whole process, because the data room already covered nearly everything. But that short list still represents real effort: the term-loan example alone would typically take a human analyst around two hours to locate, confirm, and reconcile, versus 5 minutes in Run C.

03

External access won every comparison

Across every evaluation, the run with external-data access was judged stronger, with no exceptions. Averaged across all three evaluators and both deliverables, Run C scored 13.8 of a possible 15, against Run B's 8.8 (Table 1) — the same 57% margin reported in the executive summary, but here holding in every single evaluator/task combination, not just on average.

Table 1. Blind evaluation scores, Run B (Data Room only) vs. Run C (+ external access), by evaluator and task

"Run C offers superior model precision, and more actionable handoff parameters." — Gemini, 13-Week Cash Flow

"Run C is stronger. It identifies the actual $15M quarterly covenant... and offers a more proportionate operational-defense-first recommendation." — ChatGPT, 13-Week Cash Flow

"Run C is stronger and is ready to use as-is, with its flagged assumptions kept visible." — Grok, Liquidation Valuation

04

Won by looking up facts, not by "thinking" better

Beyond Meat carries a delayed-draw term loan. Run C located a regulatory filing confirming its interest accrues rather than requiring cash payment. Without that confirmation, Run B assumed a cash-interest charge and overestimated the company's burn rate.

Discussion
01

Access alone doesn't guarantee accuracy

Access to outside data is not automatically an improvement. Foresight has to work through a specific set of checks on outside data explicitly, instead of deciding silently — does every number trace to a real source, is the output internally consistent, does a new figure conflict with anything already established. When a contradiction surfaces, Foresight keeps a record of both values and flags it, rather than quietly overwriting the old one.

What's next
AI Agents Distressed Investing External Data Fetching Blind Evaluation
Request "Foresight" Deep Dive Video