Module 07 · Risk Controls — Estimation Error and Paper Limits
Theme Concentration — Correlation and Cluster Limits
Open lesson + course map
On this lesson
Course outline
Module 1 · Market Systems and Safety — Pehle Boundaries Samjho
Module 2 · Python Bot Architecture — Ek Professional Bot Ka Skeleton
Module 3 · Market Data Pipeline — Read-Only Evidence Safely Fetch Karo
Module 4 · AI Research Engine — Extraction Se Human Review Tak
Module 5 · Strategy Research — Hypothesis Se Paper Test Tak
Module 6 · Paper Execution Engine — Synthetic Fills Only
Module 7 · Risk Controls — Estimation Error and Paper Limits
Module 8 · Database and Monitoring — Audit Logging and Model Evaluation
Module 9 · Deploying the Research Service — Read-Only and Measured
Ten paper positions can represent one underlying assumption. Markets about the same election, policy decision, team, macro release, or company event may fail together. Counting them as independent understates concentration and exaggerates sample size.
Create a versioned theme taxonomy before evaluation. Each market receives zero or more theme IDs plus an UNCLASSIFIED flag. Assignments come from deterministic metadata and human review; an optional model can propose candidates but cannot finalize them. Preserve assignment evidence and changes.
Limits operate on fictional points and counts: maximum points per theme, maximum unresolved cases per theme, and maximum share of total paper exposure. UNCLASSIFIED uses the most conservative bucket rather than bypassing controls. A case may belong to multiple themes; calculate both full-count and fractional attribution reports to show sensitivity.
Correlation estimates from small binary samples are unstable. Use them descriptively, with observation counts and time periods. Do not optimize limits from the same data used to claim performance. Stress a cluster by resolving all members adversely and by delaying their resolution together.
The dashboard should show top themes, overlapping memberships, unresolved concentration, blocked decisions, and taxonomy version. When a taxonomy update changes history, regenerate a new view while preserving the prior report hash.
Cluster controls also improve evaluation. Split related event families together across train/test windows so near-duplicate questions do not leak. Report effective event-family count alongside raw market count.
Taxonomy quality needs evaluation too. Blindly relabel a sample with a second reviewer and record agreement, ambiguous cases, and missing theme definitions. When reviewers disagree, keep all candidate themes for conservative limit checks until resolved. Do not use a model-generated embedding distance as proof of independence; it is only a candidate grouping feature with its own version and error review.
Create a theme-change impact report before accepting any taxonomy update. It lists positions newly blocked or unblocked, historical metrics affected, and event families moved across evaluation folds. The previous taxonomy remains reproducible from its manifest.
// pakistan_angle
Pakistan Angle
Pakistan political and economic questions may cluster around a small number of institutions or events. A large row count can therefore represent little independent evidence. Name the concentration instead of claiming broad local coverage, and keep the analysis paper-only.
// hands_on
Hands-On Exercise
Label twenty synthetic markets across six overlapping themes. Apply fixed limits, then stress the largest cluster. Compare raw count with event-family count. Move one ambiguous case between themes in a new taxonomy version and produce a report diff.
// completion_rubric
Completion Rubric
5 checks — tick as you verify
// sources
Sources
3 official sources — check every claim yourself