Rederivation RL round 1: blind agents rederive 5 of 9 influencer calls from our own data
Rederivation RL round 1: blind agents rederive 5 of 9 influencer calls from our own data
Question: Rederivation RL round 1: blind agents rederive 5 of 9 influencer calls from our own data Verdict: open
What we're asking
First round of the rederivation-RL loop: saved influencer insights used as a
HELD-OUT test set; 30 blind read-only subagents derived insights from
no-lookahead as-of data packs (asof-pack --verify: 0 leaks on all 10
packs); 10 grader agents scored against the held-out claims; every grade
human-reviewed. Full protocol, registry, per-case grades and derivations:
research/projects/autonomous-research/rederivation-rl/ (README, cases.json,
SCOREBOARD.md, rounds/round-01/).
Verdict: assign-follow-up
weightedScore 0.722 — 5 HIT / 3 PARTIAL / 1 MISS over 9 gradable cases
(the macro sentinel graded MISS with zero leaks, as designed).
Headline results (grades in rounds/round-01/grades/*.json, receipts in the
packs, all pack numbers derived from data/stocks/*/ohlc/12m.json bars <= T
and statements with filing_date <= T):
- Apollo Mag7 FCF-collapse call (06-28): rederived 3/3 — every agent's rank-1 was AMZN capex 122.9% of TTM OCF with FCF -25.4B vs +23.8B year-ago. This was a MISS in the 07-03 mechanical benchmark; the capex-squeeze detector + EDGAR fallback closed it same-day.
- CRWD (05-04): 3/3, with a DEEPER mechanism than the source (margin inflection to first profitable quarter, not just the price run).
- MS humanoid chokepoint (06-27): one agent derived "compact gearbox reducers = the chokepoint; Harmonic Drive ~60% supply" physics-first, blind — the held-out Morgan Stanley answer.
- AAOI (05-01): found rank-1 out of ~650 blind symbols by one agent.
- LITE beat-streak (05-04): MISS 0/3 — the load-bearing signal is paid consensus data we deliberately don't carry; free forward path is guidance-grade's own accumulated beat history.
- Systematic finding: blind agents skew bearish/risk-framed on bullish-call anomalies (SMCI, OUST graded PARTIAL on tilt, not detection).
Follow-ups filed
psRatioshare-count poisoning (in-band garbage from provider shares; false "deep value" novel insights) ->ops/tasks/TASKS-ENGINE.md#rederivation-rl #bug.- Pack
longTermblock for monthly-timeframe structure (MDB case) -> TASKS-ENGINE #rederivation-rl. - derivation-prompt v2 proposed (bull+bear framing of the largest anomaly; mandatory extreme-outlier slot) -> README, to be the single round-2 policy change.
Round 2: same cases (fixed MISSes as targets, HITs as regressions) + ~5 fresh from the curated candidate list, after the two engine fixes land.
What we found
First round of the rederivation-RL loop: saved influencer insights used as a
HELD-OUT test set; 30 blind read-only subagents derived insights from
no-lookahead as-of data packs (asof-pack --verify: 0 leaks on all 10
packs); 10 grader agents scored against the held-out claims; every grade
human-reviewed. Full protocol, registry, per-case grades and derivations:
research/projects/autonomous-research/rederivation-rl/ (README, cases.json,
SCOREBOARD.md, rounds/round-01/).
Verdict: assign-follow-up
weightedScore 0.722 — 5 HIT / 3 PARTIAL / 1 MISS over 9 gradable cases
(the macro sentinel graded MISS with zero leaks, as designed).
Headline results (grades in rounds/round-01/grades/*.json, receipts in the
packs, all pack numbers derived from data/stocks/*/ohlc/12m.json bars <= T
and statements with filing_date <= T):
- Apollo Mag7 FCF-collapse call (06-28): rederived 3/3 — every agent's rank-1 was AMZN capex 122.9% of TTM OCF with FCF -25.4B vs +23.8B year-ago. This was a MISS in the 07-03 mechanical benchmark; the capex-squeeze detector + EDGAR fallback closed it same-day.
- CRWD (05-04): 3/3, with a DEEPER mechanism than the source (margin inflection to first profitable quarter, not just the price run).
- MS humanoid chokepoint (06-27): one agent derived "compact gearbox reducers = the chokepoint; Harmonic Drive ~60% supply" physics-first, blind — the held-out Morgan Stanley answer.
- AAOI (05-01): found rank-1 out of ~650 blind symbols by one agent.
- LITE beat-streak (05-04): MISS 0/3 — the load-bearing signal is paid consensus data we deliberately don't carry; free forward path is guidance-grade's own accumulated beat history.
- Systematic finding: blind agents skew bearish/risk-framed on bullish-call anomalies (SMCI, OUST graded PARTIAL on tilt, not detection).
Follow-ups filed
psRatioshare-count poisoning (in-band garbage from provider shares; false "deep value" novel insights) ->ops/tasks/TASKS-ENGINE.md#rederivation-rl #bug.- Pack
longTermblock for monthly-timeframe structure (MDB case) -> TASKS-ENGINE #rederivation-rl. - derivation-prompt v2 proposed (bull+bear framing of the largest anomaly; mandatory extreme-outlier slot) -> README, to be the single round-2 policy change.
Round 2: same cases (fixed MISSes as targets, HITs as regressions) + ~5 fresh from the curated candidate list, after the two engine fixes land.
Verdict + reasoning
First round of the rederivation-RL loop: saved influencer insights used as a
HELD-OUT test set; 30 blind read-only subagents derived insights from
no-lookahead as-of data packs (asof-pack --verify: 0 leaks on all 10
packs); 10 grader agents scored against the held-out claims; every grade
human-reviewed. Full protocol, registry, per-case grades and derivations:
research/projects/autonomous-research/rederivation-rl/ (README, cases.json,
SCOREBOARD.md, rounds/round-01/).
Verdict: assign-follow-up
weightedScore 0.722 — 5 HIT / 3 PARTIAL / 1 MISS over 9 gradable cases
(the macro sentinel graded MISS with zero leaks, as designed).
Headline results (grades in rounds/round-01/grades/*.json, receipts in the
packs, all pack numbers derived from data/stocks/*/ohlc/12m.json bars <= T
and statements with filing_date <= T):
- Apollo Mag7 FCF-collapse call (06-28): rederived 3/3 — every agent's rank-1 was AMZN capex 122.9% of TTM OCF with FCF -25.4B vs +23.8B year-ago. This was a MISS in the 07-03 mechanical benchmark; the capex-squeeze detector + EDGAR fallback closed it same-day.
- CRWD (05-04): 3/3, with a DEEPER mechanism than the source (margin inflection to first profitable quarter, not just the price run).
- MS humanoid chokepoint (06-27): one agent derived "compact gearbox reducers = the chokepoint; Harmonic Drive ~60% supply" physics-first, blind — the held-out Morgan Stanley answer.
- AAOI (05-01): found rank-1 out of ~650 blind symbols by one agent.
- LITE beat-streak (05-04): MISS 0/3 — the load-bearing signal is paid consensus data we deliberately don't carry; free forward path is guidance-grade's own accumulated beat history.
- Systematic finding: blind agents skew bearish/risk-framed on bullish-call anomalies (SMCI, OUST graded PARTIAL on tilt, not detection).
Follow-ups filed
psRatioshare-count poisoning (in-band garbage from provider shares; false "deep value" novel insights) ->ops/tasks/TASKS-ENGINE.md#rederivation-rl #bug.- Pack
longTermblock for monthly-timeframe structure (MDB case) -> TASKS-ENGINE #rederivation-rl. - derivation-prompt v2 proposed (bull+bear framing of the largest anomaly; mandatory extreme-outlier slot) -> README, to be the single round-2 policy change.
Round 2: same cases (fixed MISSes as targets, HITs as regressions) + ~5 fresh from the curated candidate list, after the two engine fixes land.
Round 2 addendum (same day): the RL loop closed once
The three failures were refired with FREE-data pack fixes only (same
derivation prompt — data was the single variable): earnings-date estimates
(prior filing + 91d; LITE's estimate landed on the exact actual print date),
10-K Item 1 business blurbs, multi-year longTerm blocks, guidance backfill +
beat-history, SC 13D/13G stake watching, psRatio implied-shares repair.
Result: LITE MISS -> HIT (2/3 agents composed "long into the May-6 print" blind), OUST PARTIAL -> HIT (profile block flipped the read to "Lidar/ Physical AI play"), MDB stays PARTIAL (mechanism now derivable; MDB lost the attention contest to objectively louder multi-year setups like OKTA's 63-month base breakout — which the agents derived instead, blind).
Deck weightedScore: 0.722 -> 0.889 in one update cycle, zero paid data.
Round detail: research/projects/autonomous-research/rederivation-rl/rounds/round-02/ROUND.md.
Related
9 eventsNo direct external sources are attached to this read.