In Test 1, Foresight outperformed human experts using the same data room documents. In Test 2, once allowed to fetch necessary outside facts, Foresight's run with external access won every single blind comparison, scoring 57% higher on average (13.8 vs. 8.8 out of 15).
- Foresight built two modeling versions the same way at the core — the same scenario architecture and overall model design — regardless of whether it could access outside data.
- Test 2 is designed to autonomously fetch outside objective facts beyond the data room, avoiding subjective data (e.g., estimates, opinions), keeping the AI agent focused rather than drifting off-topic in a sea of open data.
- This fetching approach took under 5 minutes and consumed under 10K additional tokens, costing under $1. Total cost came in under $5 (vs. $4 in Test 1).
Takeaway: For any executive deciding whether external data access is worth enabling, this is a concrete answer: external fetch wins consistently, closing real factual gaps in a fraction of the time and cost a person would need to do the same.
This test addresses a question about AI's computing power to pull external data: when an AI agent is allowed to fetch external data in addition to a closed data room, does the resulting financial analysis improve, and by how much, and in what specific way?
The case remains Beyond Meat, Inc., the same distressed-investing opportunity analyzed in Test 1. This time, the comparison is not human versus AI — it's Foresight against itself, run under two different data-access rules on an otherwise identical task.
- Reused in full from Test 1: the same Foresight output, built from the closed data room with no external fetch permitted.
- A fresh Foresight run on the identical task, plus a narrow, tightly limited ability to look things up outside it.
- Only permitted for a hard, checkable fact (e.g., interest rate, covenant threshold, maturity date), never a forecast or opinion.
- Every lookup recorded with its source and date, checked against the data room so a new figure could never quietly replace an existing one.
- ~5 minutes, ~10K additional tokens (under $1).
- Scoring followed the same blind protocol used in Test 1: source-neutral labels, so no evaluator knew which run had external access.
Foresight built the same model twice
In Runs B and C, Foresight built both modeling versions the same way at the core: same scenario structure, same overall shape. Most checklist items matched between the two runs. Where they diverged was in execution details, not design.
Minimal external facts needed, but costly for a human to find
Run C added up to just three external lookups the whole process, because the data room already covered nearly everything. But that short list still represents real effort: the term-loan example alone would typically take a human analyst around two hours to locate, confirm, and reconcile, versus 5 minutes in Run C.
External access won every comparison
Across every evaluation, the run with external-data access was judged stronger, with no exceptions. Averaged across all three evaluators and both deliverables, Run C scored 13.8 of a possible 15, against Run B's 8.8 (Table 1) — the same 57% margin reported in the executive summary, but here holding in every single evaluator/task combination, not just on average.
"Run C offers superior model precision, and more actionable handoff parameters." — Gemini, 13-Week Cash Flow
"Run C is stronger. It identifies the actual $15M quarterly covenant... and offers a more proportionate operational-defense-first recommendation." — ChatGPT, 13-Week Cash Flow
"Run C is stronger and is ready to use as-is, with its flagged assumptions kept visible." — Grok, Liquidation Valuation
Won by looking up facts, not by "thinking" better
Beyond Meat carries a delayed-draw term loan. Run C located a regulatory filing confirming its interest accrues rather than requiring cash payment. Without that confirmation, Run B assumed a cash-interest charge and overestimated the company's burn rate.
Access alone doesn't guarantee accuracy
Access to outside data is not automatically an improvement. Foresight has to work through a specific set of checks on outside data explicitly, instead of deciding silently — does every number trace to a real source, is the output internally consistent, does a new figure conflict with anything already established. When a contradiction surfaces, Foresight keeps a record of both values and flags it, rather than quietly overwriting the old one.
- Test 3 lets the agent pull from the open internet, both objective and subjective data, to see whether that changes analysis quality.
- Test 4: In Run E, allow Claude Cowork to fetch any open data, without access to the pre-built data room. The prompt would be "Analyze distressed investment opportunity for Beyond Meat and propose a strategy." None of Foresight's prompts will be used. Verify whether this changes the quality of its analysis vs. Run D-1.