Financebotresearch desk研究台

Playbook › Playbook

fin.py --screen returns a fixed 25 rows, so "I screened the market" is really "I read 25 rows"

slow 2026-08-12

Claim

.mcp/fin.py defined def screen(name, count=25) and the argument parser called it as screen(a) with no count argument, with no CLI flag to raise it. Every run returned 25 rows under a — 25 hits header that looks like a count of matches but is a count of rows printed.

How much was hidden

Measured the same day by re-running each screen at :250:

screen rows seen true matches seen
undervalued_growth_stocks 25 238 11%
evergreen 25 156 16%
yield-tomorrow 25 130 19%
undervalued_large_caps 25 97 26%
yield-today 25 55 45%
rerating 25 45 55%
total 150 721 21%

Why it matters

  1. False exhaustiveness. Six screens feels like market coverage. It was ~21% of the matched universe, and less than that in unique names — undervalued_growth_stocks repeated roughly 60% of undervalued_large_caps.

  2. Sort order silently selects the answer, and it selects against the interesting tail. The sharpest possible demonstration: rerating matched 45 names and BR sat at row ~35, below the cut. The screen designed to find de-rated quality did match the benchmark name and then hid it. Any name found by such a screen was found despite the cap, not because of it.

How to apply

  • Pass an explicit count: --screen rerating:100. A bare name still means 25.
  • Read the returned count as a real match count only when it comes back below the requested size. rerating:100 — 45 hits is exhaustive; evergreen:100 — 100 hits is still truncated.
  • Even uncapped, screen output stays nominations only — TTM data, no multi-year history, per the Tool Hierarchy. The cap was a second, independent reason for the same rule.
  • Never write "nothing in the market beats X" off screen output. Write what the rows support.

History

  • 2026-08-12 — found while answering "are these the only ones that surfaced?" after six screens run against the BR [8.5] bar. Every screen returning exactly 25 was the tell; reading fin.py confirmed a hard-coded cap. Fixed the same day — screen(a) became a name:count partition, so --screen rerating:100 works and a bare name keeps the 25 default. Measuring the true counts afterwards showed 79% of matches had been invisible, including BR.