Rederivation RL round 1: blind agents rederive 5 of 9 influencer calls from our own data

Investigation

Rederivation RL round 1: blind agents rederive 5 of 9 influencer calls from our own data

Question: Rederivation RL round 1: blind agents rederive 5 of 9 influencer calls from our own data Verdict: open

What we're asking

First round of the rederivation-RL loop: saved influencer insights used as a HELD-OUT test set; 30 blind read-only subagents derived insights from no-lookahead as-of data packs (asof-pack --verify: 0 leaks on all 10 packs); 10 grader agents scored against the held-out claims; every grade human-reviewed. Full protocol, registry, per-case grades and derivations: research/projects/autonomous-research/rederivation-rl/ (README, cases.json, SCOREBOARD.md, rounds/round-01/).

Verdict: assign-follow-up

weightedScore 0.722 — 5 HIT / 3 PARTIAL / 1 MISS over 9 gradable cases (the macro sentinel graded MISS with zero leaks, as designed).

Headline results (grades in rounds/round-01/grades/*.json, receipts in the packs, all pack numbers derived from data/stocks/*/ohlc/12m.json bars <= T and statements with filing_date <= T):

  • Apollo Mag7 FCF-collapse call (06-28): rederived 3/3 — every agent's rank-1 was AMZN capex 122.9% of TTM OCF with FCF -25.4B vs +23.8B year-ago. This was a MISS in the 07-03 mechanical benchmark; the capex-squeeze detector + EDGAR fallback closed it same-day.
  • CRWD (05-04): 3/3, with a DEEPER mechanism than the source (margin inflection to first profitable quarter, not just the price run).
  • MS humanoid chokepoint (06-27): one agent derived "compact gearbox reducers = the chokepoint; Harmonic Drive ~60% supply" physics-first, blind — the held-out Morgan Stanley answer.
  • AAOI (05-01): found rank-1 out of ~650 blind symbols by one agent.
  • LITE beat-streak (05-04): MISS 0/3 — the load-bearing signal is paid consensus data we deliberately don't carry; free forward path is guidance-grade's own accumulated beat history.
  • Systematic finding: blind agents skew bearish/risk-framed on bullish-call anomalies (SMCI, OUST graded PARTIAL on tilt, not detection).

Follow-ups filed

  • psRatio share-count poisoning (in-band garbage from provider shares; false "deep value" novel insights) -> ops/tasks/TASKS-ENGINE.md #rederivation-rl #bug.
  • Pack longTerm block for monthly-timeframe structure (MDB case) -> TASKS-ENGINE #rederivation-rl.
  • derivation-prompt v2 proposed (bull+bear framing of the largest anomaly; mandatory extreme-outlier slot) -> README, to be the single round-2 policy change.

Round 2: same cases (fixed MISSes as targets, HITs as regressions) + ~5 fresh from the curated candidate list, after the two engine fixes land.

What we found

First round of the rederivation-RL loop: saved influencer insights used as a HELD-OUT test set; 30 blind read-only subagents derived insights from no-lookahead as-of data packs (asof-pack --verify: 0 leaks on all 10 packs); 10 grader agents scored against the held-out claims; every grade human-reviewed. Full protocol, registry, per-case grades and derivations: research/projects/autonomous-research/rederivation-rl/ (README, cases.json, SCOREBOARD.md, rounds/round-01/).

Verdict: assign-follow-up

weightedScore 0.722 — 5 HIT / 3 PARTIAL / 1 MISS over 9 gradable cases (the macro sentinel graded MISS with zero leaks, as designed).

Headline results (grades in rounds/round-01/grades/*.json, receipts in the packs, all pack numbers derived from data/stocks/*/ohlc/12m.json bars <= T and statements with filing_date <= T):

  • Apollo Mag7 FCF-collapse call (06-28): rederived 3/3 — every agent's rank-1 was AMZN capex 122.9% of TTM OCF with FCF -25.4B vs +23.8B year-ago. This was a MISS in the 07-03 mechanical benchmark; the capex-squeeze detector + EDGAR fallback closed it same-day.
  • CRWD (05-04): 3/3, with a DEEPER mechanism than the source (margin inflection to first profitable quarter, not just the price run).
  • MS humanoid chokepoint (06-27): one agent derived "compact gearbox reducers = the chokepoint; Harmonic Drive ~60% supply" physics-first, blind — the held-out Morgan Stanley answer.
  • AAOI (05-01): found rank-1 out of ~650 blind symbols by one agent.
  • LITE beat-streak (05-04): MISS 0/3 — the load-bearing signal is paid consensus data we deliberately don't carry; free forward path is guidance-grade's own accumulated beat history.
  • Systematic finding: blind agents skew bearish/risk-framed on bullish-call anomalies (SMCI, OUST graded PARTIAL on tilt, not detection).

Follow-ups filed

  • psRatio share-count poisoning (in-band garbage from provider shares; false "deep value" novel insights) -> ops/tasks/TASKS-ENGINE.md #rederivation-rl #bug.
  • Pack longTerm block for monthly-timeframe structure (MDB case) -> TASKS-ENGINE #rederivation-rl.
  • derivation-prompt v2 proposed (bull+bear framing of the largest anomaly; mandatory extreme-outlier slot) -> README, to be the single round-2 policy change.

Round 2: same cases (fixed MISSes as targets, HITs as regressions) + ~5 fresh from the curated candidate list, after the two engine fixes land.

Verdict + reasoning

First round of the rederivation-RL loop: saved influencer insights used as a HELD-OUT test set; 30 blind read-only subagents derived insights from no-lookahead as-of data packs (asof-pack --verify: 0 leaks on all 10 packs); 10 grader agents scored against the held-out claims; every grade human-reviewed. Full protocol, registry, per-case grades and derivations: research/projects/autonomous-research/rederivation-rl/ (README, cases.json, SCOREBOARD.md, rounds/round-01/).

Verdict: assign-follow-up

weightedScore 0.722 — 5 HIT / 3 PARTIAL / 1 MISS over 9 gradable cases (the macro sentinel graded MISS with zero leaks, as designed).

Headline results (grades in rounds/round-01/grades/*.json, receipts in the packs, all pack numbers derived from data/stocks/*/ohlc/12m.json bars <= T and statements with filing_date <= T):

  • Apollo Mag7 FCF-collapse call (06-28): rederived 3/3 — every agent's rank-1 was AMZN capex 122.9% of TTM OCF with FCF -25.4B vs +23.8B year-ago. This was a MISS in the 07-03 mechanical benchmark; the capex-squeeze detector + EDGAR fallback closed it same-day.
  • CRWD (05-04): 3/3, with a DEEPER mechanism than the source (margin inflection to first profitable quarter, not just the price run).
  • MS humanoid chokepoint (06-27): one agent derived "compact gearbox reducers = the chokepoint; Harmonic Drive ~60% supply" physics-first, blind — the held-out Morgan Stanley answer.
  • AAOI (05-01): found rank-1 out of ~650 blind symbols by one agent.
  • LITE beat-streak (05-04): MISS 0/3 — the load-bearing signal is paid consensus data we deliberately don't carry; free forward path is guidance-grade's own accumulated beat history.
  • Systematic finding: blind agents skew bearish/risk-framed on bullish-call anomalies (SMCI, OUST graded PARTIAL on tilt, not detection).

Follow-ups filed

  • psRatio share-count poisoning (in-band garbage from provider shares; false "deep value" novel insights) -> ops/tasks/TASKS-ENGINE.md #rederivation-rl #bug.
  • Pack longTerm block for monthly-timeframe structure (MDB case) -> TASKS-ENGINE #rederivation-rl.
  • derivation-prompt v2 proposed (bull+bear framing of the largest anomaly; mandatory extreme-outlier slot) -> README, to be the single round-2 policy change.

Round 2: same cases (fixed MISSes as targets, HITs as regressions) + ~5 fresh from the curated candidate list, after the two engine fixes land.

Round 2 addendum (same day): the RL loop closed once

The three failures were refired with FREE-data pack fixes only (same derivation prompt — data was the single variable): earnings-date estimates (prior filing + 91d; LITE's estimate landed on the exact actual print date), 10-K Item 1 business blurbs, multi-year longTerm blocks, guidance backfill + beat-history, SC 13D/13G stake watching, psRatio implied-shares repair.

Result: LITE MISS -> HIT (2/3 agents composed "long into the May-6 print" blind), OUST PARTIAL -> HIT (profile block flipped the read to "Lidar/ Physical AI play"), MDB stays PARTIAL (mechanism now derivable; MDB lost the attention contest to objectively louder multi-year setups like OKTA's 63-month base breakout — which the agents derived instead, blind).

Deck weightedScore: 0.722 -> 0.889 in one update cycle, zero paid data. Round detail: research/projects/autonomous-research/rederivation-rl/rounds/round-02/ROUND.md.

9 events

No direct external sources are attached to this read.