Back to Foresight
Test 3 · Aug 2026

Distressed Investing Test 3

Unrestricted Open Data Destroys AI Agent's Outputs

Built Aug 2026
AI agent Foresight — Claude Opus 4.8, High Thinking
AI Evaluators ChatGPT 5.5 + High Thinking, Grok Expert (Gemini unable to complete)
Published on Towards Finance
Executive Summary

In Test 1, Foresight v1 — working from a closed data room — outperformed senior human experts, and in Test 2, letting it fetch only the necessary missing facts improved its output further. In Test 3, once Foresight fetches open data entirely on its own, without a data room at all, output quality drops significantly. This run scores roughly 44% lower on average than the prior two runs, though it is still 7.5% higher than senior human experts.

Takeaway: Open-data fetching is still directionally reliable despite the drop in accuracy — feasible and cost-effective as a first-pass screen across a large set of opportunities, at a fraction of the cost of senior human experts. The prompt still needs refinement to avoid missing key data already sitting in its own sources. For any executive deciding whether unrestricted open-data access is worth enabling, this is a concrete answer: it still beats senior human experts, and it opens the door to screening opportunities directly from real-time market signals.

This test addresses a question about AI's autonomous judgment over which data to trust: when an agent collects data entirely from open sources on its own, rather than from a closed data room, does the financial analysis drift — and if so, by how much, and at what cost?

Test 3: Beyond Meat, distressed investing (Aug 2026)

The case remains Beyond Meat, Inc., the same distressed-investing opportunity analyzed in Test 1 and Test 2. The comparison is again Foresight against itself — now run under unrestricted, open-data fetching rules on an otherwise identical task, subject to the same Dec 8, 2025 cutoff used in the earlier tests.

How the test was set up
Run A
Human · Data Room only
  • Completed by the four-person Stern EMBA team.
Run B
AI · Data Room only
  • ~1 hours, 156K output tokens (~$4).
    • Reused in full from Test 1: the Foresight output built from the closed data room, no external fetch.
    Run C
    AI · Data Room + restricted fetch
    • Reused in full from Test 2: the closed data room plus permission to fetch specific missing facts only.
    • ~1 hours, 166K output tokens (~$5).
    Run D
    AI · Open data + No Data Room
    • A fresh Foresight run on the identical task, with no data room at all and full ability to look up anything from the internet (same Dec 8, 2025 cutoff).
    • Run in Claude Cowork so each lookup's full context could be saved and organized.
    • ~1.5 hours, 1,278K output tokens (~$32).
    Blind Evaluation
    Scoring method
    Findings
    01

    On-topic sources, but press-release-heavy

    Run D's self-collected set was all on-topic for distressed investing and included the primary financial filing needed for modeling. But it skewed to news: ten of twelve items were company-specific news against a single financial statement, where the closed data room was anchored by SEC filings.

    Table 1. Type of data source: Run D collected from open data versus the closed data room used by Runs A, B, and C. Run D's set was press-release-heavy (10 of 12 company-specific news items, 1 financial statement) while the data room was anchored by SEC filings.
    02

    No AI hallucination

    Even though Run D collected its own data and built the model from it, both evaluators explicitly checked and confirmed that sampled figures traced back to real documents, with no fabricated sources. The prompts' built-in safeguards — explicit reasoning, a self-assessment pass, and mandatory source traceability — held.

    03

    Two post-quarter figures never made it into the model

    Run D missed Beyond Meat's $148.7M at-the-market (ATM) equity raise in the 13-week cash flow, and used pre-exchange debt instead of the post-exchange figure in the liquidation valuation. Both emerged after Q3 2025. One was dropped at data capture — a cash-positive raise ran counter to the distress thesis, and no prompt forced the agent to reconcile it; the other was captured only as fragments the agent never summed into a single number.

    04

    Open-data access lost to both prior runs

    Run D-1 slightly beat Run A, where its cleaner model architecture outweighed its shortfall on the two missed figures. Against Run B and Run C — both anchored by the data room — the open-data run was judged weaker in every comparison.

    "Run D-1 delivers a complete, auditable, multi-scenario model with graded critical assumptions, explicit working-capital stress, self-validation checks, and a clear decision handoff." — Grok, 13-Week Cash Flow (Run D-1 vs. Run A)

    "Run A better captures the post-Q3 financing/capital-structure change and produces security-level recoveries, while Run D-1 leaves its headline denominator at pre-exchange funded debt and omits the documented ATM inflow." — ChatGPT, Liquidation Valuation (Run D-1 vs. Run A)

    Table 2. Blind evaluation scores, Run A (human) vs. Run D-1 (open data), scored by ChatGPT and Grok on the 13-Week Cash Flow and Liquidation Valuation deliverables. Average of 15: 10.0 / 10.75.

    "Run B is stronger because it includes the documented ATM proceeds and therefore gives the more defensible direction for near-term cash. Run D-1 has cleaner PIK debt-service treatment and more direct primary-source labeling, but omitting the $148.7 million ATM materially understates cash and overwhelms those advantages." — ChatGPT, 13-Week Cash Flow (Run D-1 vs. Run B)

    Table 3. Blind evaluation scores, Run B (data room only) vs. Run D-1 (open data), scored by ChatGPT and Grok on the 13-Week Cash Flow and Liquidation Valuation deliverables. Average of 15: 12.75 / 7.75.

    "Run C is ready to use as-is for hand-off to the Priority-3 debt waterfall. Run D-1 should not be used for any Dec-8 decision without a complete rebuild of the valuation-date balance sheet and the post-exchange claim stack." — Grok, Liquidation Valuation (Run D-1 vs. Run C)

    Table 4. Blind evaluation scores, Run C (data room + restricted fetch) vs. Run D-1 (open data), scored by ChatGPT and Grok on the 13-Week Cash Flow and Liquidation Valuation deliverables. Average of 15: 14.25 / 7.5.
    Discussion
    01

    Root-cause analysis still needs senior human experts

    AI cannot investigate the root cause of missing data on its own; it stays at the surface and still depends on senior human experts to challenge its answers. But Cowork — where each run produces 20+ generated files — speeds up the file-by-file comparison, so an expert can pinpoint the root cause within about an hour.

    02

    Open-data fetch is still unstable

    A second collection round gathered less than the first — pulling only the Q3 2025 earnings-call transcript rather than the 10-Q — which shifts the analytical plan and how the model is built. Stabilizing the fetch is targeted for development under Foresight v2.

    Table 5. Type of data source, Run D-1 (first collection round) vs. Run D-2 (second round). The second round gathered fewer sources, pulling the Q3 2025 earnings-call transcript instead of the 10-Q filing.
    What's next
    AI Agents Distressed Investing Open Data Fetching Blind Evaluation
    Request "Foresight" Deep Dive Video