How often are the agents right?
Every day we save what the agents said, wait, and then check it against what actually happened. No cherry-picking — every call counts.
Live Track Record
5-day horizonEach daily signal is recorded and scored 5 trading days later: a bullish call is “correct” if the price is higher, a bearish call if lower. This is the live record — it grows every day and cannot be edited. Collecting since 7/13/2026.
These four numbers are the deterministic advisor’s alone — the engine behind today’s signals. The retired editorial baseline is kept in the record for honesty and reported separately below; the two are never added together. Across every stream, 99 calls have been evaluated and 71 are pending.
Which engine produced the calls
Every row in the record is labeled with the stack that produced it, and the two streams are never blended into one number. “Deterministic advisor” is the live quant engine (measured candles, ADX, volatility, calibrated confidence) that generates today’s signals; “editorial baseline” is the fallback it replaced, kept here because its history is part of the honest record.
| Signal engine | Recorded | Evaluated | Pending | Accuracy | 95% range | Avg move |
|---|---|---|---|---|---|---|
| Deterministic advisor (the quant engine) | 43 | 31 | 12 | 42% | 26–59% | +0.255% |
| Editorial baseline (pre-engine fallback) | 127 | 68 | 59 | 57% | 46–68% | +0.403% |
| Market | Calls recorded | Evaluated | Correct | Accuracy |
|---|---|---|---|---|
| XAUUSD | 14 | 9 | 6 | 67% |
| AAPL | 14 | 9 | 4 | 44% |
| DXY | 14 | 9 | 6 | 67% |
| USOIL | 14 | 9 | 0 | 0% |
| ETHUSD | 14 | 9 | 5 | 56% |
| TSLA | 12 | 7 | 7 | 100% |
| BTCUSD | 12 | 7 | 4 | 57% |
| SPY | 12 | 7 | 1 | 14% |
| QQQ | 12 | 7 | 2 | 29% |
| NVDA | 12 | 7 | 4 | 57% |
| SOLUSD | 12 | 7 | 4 | 57% |
| EURUSD | 7 | 3 | 3 | 100% |
| USDCAD | 7 | 3 | 3 | 100% |
| USDJPY | 7 | 3 | 3 | 100% |
| AUDUSD | 7 | 3 | 0 | 0% |
No accuracy claim is made below 30 evaluated signals — until then this page reports collection progress only. For the multi-year simulation of the same logic, see the admin Backtest Lab. Past accuracy does not guarantee future results.
Does the stated confidence mean anything?
calibration error 16ptsThe receipt behind every confidence number: signals are grouped by what we said (stated confidence) and scored by what happened (measured hit rate). Perfect honesty means the two columns match. Calibration is reported per engine — a calibrated stream blended with an uncalibrated one would describe neither. This table is machine-readable at /api/track-record/reliability and cannot be edited.
Deterministic advisor (the quant engine)
31 evaluated signals; calibration error 8 points. 12 pending.
| We said | Signals | Measured hit rate | 95% range | Status |
|---|---|---|---|---|
| 45–50% | 13 | — | — | collecting 13 of 30 |
| 50–55% | 18 | — | — | collecting 18 of 30 |
Editorial baseline (pre-engine fallback)
68 evaluated signals; calibration error 20 points. 59 pending.
| We said | Signals | Measured hit rate | 95% range | Status |
|---|---|---|---|---|
| 60–65% | 8 | — | — | collecting 8 of 30 |
| 65–70% | 13 | — | — | collecting 13 of 30 |
| 70–75% | 10 | — | — | collecting 10 of 30 |
| 75–80% | 12 | — | — | collecting 12 of 30 |
| 80–85% | 25 | — | — | collecting 25 of 30 |
Reliability = measured hit rate per stated-confidence bucket over the live, append-only signal record, reported per signal stream: 'advisor_live' is the deterministic quant engine, 'seed_baseline' the editorial fallback it replaced. Buckets under 30 evaluated signals are still collecting and make no claim. Past accuracy does not guarantee future results.
Can people actually read this?
measured, not assertedWe claim the product explains itself to beginners. Onboarding ends with three questions about how to read a card — what a model score is, what a stop means, and what “follow with paper money” does. These are the answers, including the wrong ones. A rate only counts as a claim at 50 answers per mode.
No one has taken the check yet. It appears at the end of onboarding.
Comprehension is a claim only at N >= 50 per mode; below that this is collection progress.