Methodology: How the Fair Line Is Built
Every venue posts a different price on the same game. Here is exactly how we turn that into one number, and when we refuse to.
What this runs over today: MLB, NFL and the rest of the season's slate. Around a dozen venues, the Kalshi exchange among them, on moneylines, pregame. That roster is small, and it is the reason several of the refusal rules below fire as often as they do.
The short version
Every venue's price has that venue's cut baked into it, so the first job is to take the cut back out. What is left is what that venue actually thinks will happen.
Then we pool those opinions. Not by averaging them, which quietly overweights the middle, but in log-odds, which is the right way to average a probability. Venues that are tight and deep count for more than venues that are wide and thin, and venues that copy each other get discounted so one opinion is not counted twice.
When we grade a venue's price, we pool everyone else and compare against that. A price measured against a benchmark it helped set always looks fair, which is the most common way to accidentally invent an edge that is not there.
And when the market is too thin or the venues disagree too much, we say so instead of printing a confident number. Every fair line ships with how many venues stood behind it.
That is the whole thing. The rest of this page is the same nine steps with the arithmetic attached, so a number you disagree with can be reproduced by hand or against the published library. These are not the only defensible choices. They are the ones we made, written down.
The pipeline, in order
- Extract one price per venue. Microprice on an order book, posted two-way moneyline on a book.
- Remove the margin. De-vig each venue independently: proportional, power or Shin.
- Correct for favorite–longshot bias. The power exponent shades longshots down.
- Pool in log-odds. A logarithmic opinion pool, not an average of probabilities.
- Weight each venue. A cold-start tier plus two measured quality terms today; fitted calibration weights as the record accumulates.
- Discount correlated venues. Arbitrage-linked venues are not independent opinions.
- Grade leave-one-out. A venue is never measured against a benchmark it helped set.
- Gate on quality, and publish the count. Too few venues, or too much disagreement, and that is shown rather than smoothed away.
- Score the evidence behind the line. Agreement, the effective number of independent venues, and whether a trusted one is among them, combined into the published confidence figure (formula and constants in step 9).
Steps 1–4 are ordinary published statistics, documented in full below. Step 5’s fitted multipliers and step 6’s measured correlation are the one thing this site will keep private once they exist. See what is not published. The cold-start weighting running in their place is published in full, formula and constants, in step 5.
Step 1: The price taken from each venue
An exchange quotes a two-sided book, not one number. The naive summary is the mid, halfway between best bid and best ask, and it throws away the most informative thing on the screen: how much size is resting on each side. The microprice keeps it.
where:
- : best bid, as a probability in [0, 1] (a 62¢ contract is 0.62)
- : best ask, same units
- : size resting at the best bid, in contracts
- : size resting at the best ask, in contracts
Note the crossing of the subscripts: the bid is weighted by the size on the ask side. That is what makes the estimator lean toward the thinner half of the book, the half that clears first, and so the direction the price is about to move.[5]
Worked example: microprice vs. mid
On a binary exchange the two half-books are reconciled first. Buying NO at is selling YES at , so the YES ask is one minus the best NO bid. If either half-book is empty, or the derived ask sits at or below the bid, the venue is not quoted at all: a crossed book is a data error, and a data error is not a price.
A sportsbook has no order book. Its posted two-way moneyline is the price, and step 2 does the work instead.
Step 2: Removing the margin
A posted two-way line implies two probabilities that sum to more than 1. The excess is the venue’s margin, the vig. De-vigging is the arithmetic that removes it.
where is the American moneyline and the vig-inclusive implied probability. Three methods split the overround differently, and they are not interchangeable.
Proportional
Divides the excess in proportion to each side’s implied probability. It is the baseline: one line of arithmetic, no parameters, and it is what most calculators mean by “no-vig.” It also assumes the margin is spread evenly across the price scale, which is empirically false.
Power
Raises both implied probabilities to a common exponent found by bisection. Because raising a number below 1 to a power greater than 1 shrinks the small one proportionally harder, the power method takes more away from the longshot than from the favorite. That matches the direction of the observed favorite–longshot bias.[3][4] This is the method behind the pooled fair line on the board.
Shin
Shin’s model treats the margin as the venue’s defence against a fraction of informed money, and solves for the that makes the fair probabilities sum to 1.[1][2] It is the right tool when the question is specifically how much of this price is insider protection: a thin market, an injury rumour, an early number.
Worked example: the same line, three methods
The spread between the three methods is retained per venue as a data-quality signal. A line the three methods agree about is a clean two-way price. A line they disagree about is usually mis-paired, mis-timed or mistyped, and it earns less weight in the pool. Run any of the three yourself on the no-vig calculator; the long walkthrough is how to de-vig odds.
Step 3: Favorite–longshot calibration
Longshots are systematically overpriced and favorites underpriced, across parimutuel pools, sportsbooks and event-contract exchanges alike.[3][4] A price of 5¢ on a binary contract wins less often than 5% of the time, and after fees the gap widens further, because the fee is a larger share of the money at risk at the tails (see fair value is not break-even).
Two consequences. The power exponent in step 2 is the mechanism that would apply this correction, and no single method is trusted with the displayed line alone: the default runs all three methods on every market and keeps the most generous grade, printing the winner beside the number (the max-of-three selection). The exponent power depends on is meant to be fitted against settled markets and has not been yet. Closing-line grades and the venue calibration ledger stay on the proportional method alone, the one a reader can reproduce by hand, so the archived record keeps a single consistent ruler. Any one method stays selectable, and the number is always published with the method that built it. And no venue is treated as an unbiased oracle: an exchange price is money at risk, which makes it informative, not calibrated.
The exponent used on the board today is the one that makes each venue’s two sides sum to 1: a per-line solve, not a fitted seasonal constant. Fitting against settled markets requires a settled-market archive; that archive is being written now and does not yet move the published line. See what runs today.
Step 4: Pooling in log-odds
Several venues, one number. Averaging their probabilities is the wrong arithmetic: probability is a bounded scale and a linear average of forecasts is provably underconfident, drifting toward the middle.[6] The pool runs in log-odds.
where is venue ’s de-vigged probability for side A, its weight, and the published fair probability. Side B is by construction, so the published line carries exactly zero vig, a property the arithmetic guarantees rather than something rounded into place.
Exponentiate and the sum becomes a product: the pooled odds ratio is a weighted geometric mean of the venues’ odds ratios. That is also why duplicate feeds are so damaging: a copied line is the same factor repeated, one opinion wearing several exponents.
Worked example: log-odds pool vs. plain average
Step 5: Weights
The end state: sharpness treated as a property of a market at a moment, not of a brand. That is not yet what runs. Today a venue starts from a hardcoded cold-start tier, and because that tier is the single largest factor in the shipped weight, it is printed here in full.
True either way: no venue can buy weight. No commercial relationships with any venue, no paid placement, no affiliate links anywhere on this site.
The weight that runs today is:
where:
- : the cold-start tier for venue , dimensionless. 2.0 for exchanges (Betfair, Betfair Exchange, Smarkets, Matchbook, Kalshi, Polymarket); 1.9 for Prophet Exchange and ProphetX; 1.8 for Pinnacle, Circa and Circa Sports; 1.6 for BookMaker and BetOnline; 1.5 for Heritage; 1.0 for every venue not on that list. Read that list as a table of constants, not a roster. Of the venues named in it, the board carries exactly one today: Kalshi. Betfair, Smarkets, Matchbook, Polymarket, Pinnacle, Circa, Prophet and Heritage are not on this board, so their numbers above have never been applied to a published line. BetOnline and Bookmaker are on the board and are still not lifted, because they sit in the offshore tail that is excluded from the pool entirely (a licensing posture, explained in step 8, not a view about their prices). So the honest summary of the tier term as it runs: it lifts Kalshi, and it lifts nothing else. What the other rows describe is what would happen if those venues were ever carried, which is a plan, and it is written down here so the plan cannot be mistaken for the product.
- : the venue’s overround on this market, as a decimal (a 3.5% hold is 0.035)
- : the step-2 disagreement between the three de-vig methods on this venue’s line, in probability
- : the reference two-way hold; : a floor so a mis-parsed near-zero-vig line cannot dominate
- : the margin exponent, deliberately small; : the de-vig-disagreement penalty
Those constants are worth reading as numbers:
- Margin spans about 0.78 to 1.19 across the holds a real book posts (5% → 0.7765, 2% → 0.9457, 0.5% → 1.1892), against a tier term spanning 1.0 to 2.0. Compressed relative to tier, but a genuine factor rather than a tie-breaker.
- Disagreement sits near 1 for a clean line (0.967 on the −150 / +130 example above) and falls to 0.67 once the three methods differ by 5 probability points. An ambiguous line is discounted, not trusted at face value.
- A tier at or above 1.5 marks a venue as the pool’s anchor: what the step-8 outlier clip is centred on, and what the confidence cap keys off.
No single venue, or cluster of duplicated feeds, is allowed more than half the total weight.
The tier is a cold-start prior, not a permanent judgement, precisely what the fitted weights are being built to replace. What matters is measured calibration: every venue’s closing price is logged and every settled market harvested, so each venue can be scored on how well its early price anticipated the close and tracked realised outcomes. Those multipliers are not yet computed and do not move the published line. See what runs today.
The target weight, once the streamed order-book lane drives the board, adds the two observable quantities the shipped formula has no access to (resting size, and time since the venue last moved the price):
where is the venue’s quoted spread in probability, the size resting at the touch in contracts, the seconds since the venue last moved this price, and the staleness time constant. Neither depth nor a time-decay term is an input to the published weights today. Staleness is a binary cut, not a decay: a venue more than 30 minutes behind the freshest comparable soft book (5 minutes for Kalshi, which is re-polled every minute) is dropped from the signals outright (step 8).
One deliberate limit on the margin term: a thin margin is not evidence of sharpness. Soft books post their tightest holds on exactly the high-volume public games they shade hardest, so a pool weighted mainly by low vig hands the loudest voice to the most shaded prices. That is why is 0.25 rather than 1: margin is compressed on purpose, not because it is uninformative.
Step 6: Correlated venues
Two prediction markets on the same event are arbitrage-linked. When one moves, the other follows, because someone is paid to make it follow. Pooling them as independent draws counts one opinion twice and manufactures confidence that is not there.
Two mechanisms are specified; one of them runs today.
Runs: near-identical lines are collapsed into a single vote before any weighting, so a monoculture of copied feeds cannot inflate the pool, and the effective venue count behind the confidence read is deflated rather than counting logos.
Does not: the residual correlation between venues that are genuinely distinct but arbitrage-linked, measured over time and used to discount their combined weight. It needs the same archive as step 5. See what runs today.
Step 7: Leave-one-out grading
This one is a correctness rule, not a refinement. When a venue’s price is graded against a benchmark, the benchmark must be recomputed without that venue.
Skip it and every venue is partly measured against itself. Its own price is inside the number it is being compared to, the two are dragged toward each other, and every edge and every closing-line-value grade comes out systematically compressed toward zero. The bias is largest exactly where it matters most: on a thin slate, where one venue is a large share of the pool.
Worked example: what self-inclusion hides
The same rule governs the closing line value grades and the venue-calibration scoring in step 5: a venue’s opening price is scored against a close it did not help set.
Where fewer than two venues remain after the exclusion, there is no leave-one-out benchmark to compute, so the full-pool number is used rather than a one-venue benchmark being invented. That fallback is not flagged in the payload beside the grade, which it should be; the venue count on the row is the signal that it may have applied.
Step 8: Quality gates, and what happens when the data is thin
The first gate is an exclusion, not a roster. Every venue that can carry a mark also helps set the fair line: one pool, for the pooled fair line, the +EV flags and the arbitrage check alike. Until 7 August 2026 there was a second, narrower list (DraftKings, FanDuel, BetMGM, Caesars, Kalshi) that alone could vote, which meant a book could be flagged as the best price on the board without being allowed to count toward the number it was measured against. That was a hardcoded venue reputation, which step 5 says not to do, and the rule it protected (no venue helps set the line it is graded against) is already step 7’s job. So it is gone.
What is still excluded, and excluded absolutely, is the offshore tail: LowVig.ag, BetOnline.ag, MyBookie, Bookmaker, BetAnySports and Everygame. They stay on the board, priced and timestamped, because a second opinion is worth reading. They never produce a flag and never vote, and that is a licensing decision rather than a view on their prices: a flag is the closest this site comes to pointing at a venue and saying “there”, and it should never point at one that is unlicensed in the US.
Every other book on the row shows its price and timestamps but is marked reference-only: it cannot vote in the benchmark and is never flagged as an edge. If fewer than two roster venues quote a market, the fallback is every book on the row rather than publishing nothing.
Beyond the roster, an exchange quote is excluded when its spread exceeds 6 cents or either half-book is crossed or empty. A wide market is still ingested and displayed with its true bid, ask and timestamp; it just cannot vote. The depth floor described in step 1 is specified but not applied: no resting-size test gates a quote today.
A venue whose price has frozen while its peers moved is demoted to reference-only: a soft book more than 30 minutes behind the freshest comparable soft book, or the exchange more than 5 minutes behind it (Kalshi is re-polled every minute, so a wider gap is a dead feed, not lag). That cut no-ops rather than leaving fewer than two voters. A frozen line is not an opportunity, it is a broken feed, and it is how a stale row turns into a phantom arbitrage.
Outliers are handled by a redescending robust weight centred on the pool’s anchor, not the crowd’s median. Centre the clip on the crowd and the sharpest price gets “corrected” as the outlier, which is exactly backwards. A venue far enough from the anchor is down-weighted continuously toward zero: a price that far out of line is almost always mis-paired, mis-timed or mistyped rather than information.
Two limits: the clip engages only at four or more distinct venues, and where the pool contains no anchor the clip falls back to the weighted median (the crowd), which is the case the confidence cap below exists for.
Then the refusal rules, all four running on the published board as of 5 August 2026:
The refusal rules
- Fewer than two qualifying venues → the board refuses. One venue is not a consensus, and dressing a single quote up as one is the failure mode this page exists to avoid. Since 5 August 2026 the published board withholds the fair line and prints the reason instead: a venue count and a label, with the prices still shown, timestamped, reference-only. A market whose analysis throws degrades to the same refusal rather than an error. The calibration ledger applied this gate already; the board now matches it.
- Dispersion above the wide threshold → the disagreement is published with the number. Every fair line ships with its venue count and an agreement label: tight, moderate or wide. A wide market is a real state of the world, and hiding it behind a confident-looking number is a lie of presentation.
- No anchor-tier venue in the pool → the confidence figure is capped at 0.45. A crowd of soft books that all copy each other is not a benchmark, however many of them there are. Note what this cap keys off today: the hardcoded tier of step 5, not a measured price-discovery record. Replacing the one with the other is the whole point of the calibration ledger.
- No usable quote → the row is blank. Not interpolated, not carried forward from the last poll.
Same instinct as the honest zero: an implemented failure state is visible rather than papered over. Where a fallback is still silent, in the roster and leave-one-out cases above, this page says so, because a methodology page that lists only the gates that work is not a methodology page.
Step 9: The confidence figure, in full
Every fair line on the board ships with a confidence percentage and an agreement word. Both are published; neither was written down until now. Here is the whole calculation.
First, what it is not. It is not how sure we are that the arithmetic is right, and it is not the chance the bet wins. It is how much evidence stands behind this particular fair line: thin when few venues quote the market or they disagree, high when many independent venues agree and one of them is trusted to lead. A confidence of 35% does not mean the maths is 35% reliable; it means this line rests on less than a well-covered market would give you.
It is the product of two factors, each in [0, 1], and then one adjustment:
confidence = exp(−σz / 0.25) × neff / (neff + 2)
- How much the venues disagree. σz is the weighted spread of the venues’ de-vigged log-odds around the pooled value, the same dispersion the agreement word is cut from. It enters as exp(−σz/0.25), so perfect agreement scores 1 and confidence falls off smoothly as the market splits: a dispersion of 0.25 costs about 63% of the figure, 0.50 about 86%. Log-odds, not probability points, so a disagreement at 95% counts for more than the same gap at 55%.
- How many independent opinions there really are. neff = (∑w)² / ∑w², the effective sample size of the weights: five venues at equal weight give 5, but five where one carries most of the weight give barely more than 1. It enters as neff/(neff+2): two effective venues score 0.5, four score 0.67, ten score 0.83, and no number of venues ever reaches 1. Copycat feeds collapse before this is computed (Step 6), so ten mirrors of one line cannot inflate it.
Then the adjustment, which is where the anchor matters:
- With an anchor-tier venue in the pool, the figure is multiplied by 0.5 + 0.5 × (anchor share of the weight). A pool led by its anchor keeps nearly all of its confidence; one where the anchor is a bystander keeps half.
- With no anchor-tier venue, the figure is capped at 0.45 instead. That is the cap referenced throughout this page: a fair line pooled entirely from soft books may be the best estimate available, and it is still not allowed to present itself as a confident one.
The result is clamped to [0, 1] and shown as a percentage. The agreement word beside it is σz alone, in three bands: tight below 0.08, moderate below 0.20, wide above. So a board can honestly read “tight agreement” and a middling confidence at the same time: that is two venues agreeing closely, which is agreement without much evidence.
What stays private is unchanged by publishing this: the formula is here, but which venues hold anchor tier and what each one weighs are not, so the anchor-share term is reproducible in form and not in value. Everything else (the dispersion, the effective count, the cap) you can recompute from the venue count and prices the board already prints.
How the EV lane is ordered
The EV lane sorts by the pooled EV figure. Two things can shelve a row without touching that figure.
First, your roster. A flag at a venue you have ticked sits above one you have not. That is delivery, not math: the dimmed row still counts, and its price still votes in the pool.
Second, the line move. The line-move shelf is on: a row whose side’s pooled fair line fell 0.5 points or more in the last 60 minutes (three hours for a game more than 14 days out, whose prices are checked hourly) sits below rows whose side held, inside those roster shelves. The move is measured with the proportional de-vig over the same venues at both clocks, with the flagged venue left out, the same way the EV is graded (step 7), and it is refused when fewer than two other venues priced the market at either clock. The shelf is off today; see what runs today.
Neither shelf changes an EV, a flag, or which cases can alert; the alert list adopts the same order so a phone leads with the case the board leads with. Every reading prints its two clocks on the row and on the game card.
Fair value is not the break-even price
These are two different numbers, and conflating them overstates every edge computed from them. Fair value is the pooled probability, with no margin in it. Break-even is the probability a position needs to clear the price plus the fee. They are reported separately.
On Kalshi the fee is closed-form and charged per trade:
where is contracts and the price in dollars. Dropping the cent rounding gives the break-even in closed form:
Worked example: a 50¢ contract
Wherever an exchange price is treated as tradeable (an arbitrage leg, a best price, an EV flag, a stake size), the number used is the executable taker cost (ask plus fee), not the fair mid. The fair mid feeds the pool; the executable cost gets graded.
Some major events instead carry a flat 0.25% maker fee, a large fraction of a sub-5¢ contract’s price; that case is handled explicitly rather than folded into the standard formula. The Kalshi fee calculator prints both numbers side by side.
The suspect band
Expected value per dollar staked, against the leave-one-out benchmark:
where is the benchmark fair probability and the net decimal profit per $1 at the posted price. A ceiling applies:
Any quote computing to 15¢ or more of edge per dollar is dropped from the flags rather than surfaced. In a liquid two-way market a fifteen-cent edge is not an edge; it is a stale price, a mis-paired line, a postponed game, or a typo. Filtering the top of the distribution costs a genuine outlier occasionally. Publishing it costs the reader real money, which is worse.
Worked example: flagged, and filtered
An arbitrage flag carries the same kind of floor. A cross worth less than $1 per $1,000 staked is not published as an arbitrage, because after two stakes and the time it takes to place them it is not one.
The honest zero
When the inputs contain no edge, the printed answer is $0. Not “marginal.” Not a small stake rounded up so the screen has something on it.
Worked example: a no-edge input
Timestamps, staleness and where the data comes from
Every quote carries two timestamps: when it was pulled, and when the venue last moved it. Both are displayed. A live streamed price and a snapshot polled three hours ago are rendered differently, because they are different kinds of fact, and a tool that shows them identically is misrepresenting the more useful one.
Sources are first-party licensed APIs only: the Kalshi exchange API for prediction-market prices, a commercial odds API for sportsbook lines. Nothing is scraped, and no resold or scraped feed is purchased. That constraint has teeth: it rules out resting the benchmark on a resold feed from a single sharp book, which is why the fair line here is a pool at all.
Every observed price is written to an append-only archive keyed on event, venue, market and timestamp, and settled outcomes are harvested from the exchange. Nothing is overwritten in place. That archive is what makes step 5’s measured weights and the closing-line-value grades possible at all.
What is not published
One thing, stated precisely so the boundary is not mistaken for evasion: the fitted per-venue calibration weights and the measured correlation structure between venues, once they exist. Neither is computed yet. The cold-start tier table and weight formula running in their place are published above in step 5, constants and all. Nothing is withheld about the estimator actually running.
What is published, on every fair line: the resulting probability and price, how many venues qualified, how much they agreed, the confidence read, the de-vig method used, and every timestamp. What is published on this page: the estimator for every step, with its formula. The technique is not the asset. The record of which venues have actually predicted closes and outcomes, accumulated one settled market at a time, is.
Nothing else is withheld. The odds conversions, all three de-vig methods, Kelly sizing and the arbitrage split are MIT-licensed and installable from PyPI, precisely so a number here can be checked without asking permission.
What runs today
Stated plainly, because a methodology page that describes an aspiration as a fact is worse than no methodology page.
- Live: per-venue de-vig (all three methods), the proportionally de-vigged log-odds pool with power and Shin selectable, duplicate-feed collapse, the anchored outlier clip, the freshness gate, leave-one-out grading on every edge and every closing-line-value grade, the executable taker-cost substitution for exchange prices, the suspect band, the arbitrage floor, the published confidence figure and agreement word (step 9, formula and constants included), the append-only tick archive, and settled-outcome harvesting.
- Live, and hardcoded rather than measured: the tier table in step 8, which sets each venue’s starting weight and which venue counts as the anchor. The only venue it lifts above 1.0 is Kalshi; every book appears nowhere in the table and takes the 1.0 default. The companion list, a fixed roster of which venues could vote at all, was removed on 7 August 2026 for exactly the reason this section exists to admit: it was a brand list doing a measurement’s job. The tier table is the last of it, it is a hand-chosen cold-start prior, it is published in full above, and it is what the fitted weights are meant to replace.
- Live since 7 August 2026: the streamed order-book lane. When the stream is up, the Kalshi quote reaching the pool is derived from the live book, with the executable ask and the top-of-book depth captured alongside it; when it is not, the scheduled snapshot stands, stamped as such. Depth is recorded but is not yet an input to the published weights; the fitted per-venue weights remain the open item, stated below.
- Live since 2 September 2026: the line-move reading, on every flagged +EV case on the live board and its game card. The shelf that sorts on it was armed by the owner the same morning, ahead of the pre-registered backtest, which reports on it monthly (
docs/decisions/sharp-money-discriminator.md). - Accumulating, not yet applied: the fitted per-venue calibration weights of step 5 and the measured correlation discount of step 6. The archive they need is being written now. Until it is deep enough to be worth trusting, the cold-start weighting is what runs, and this page will say so until that changes.
This section changes when the code does, and the date in the byline is the same string as the dateModified in this page’s structured data, so the two cannot drift apart.
Check any of it
Reproduce a number
pip install teachersbettextbook
from teachersbettextbook.devig import devig_power, overround devig_power(-150, 130) # -> (0.5840, 0.4160) overround(-150, 130) # -> 0.03478
Same inputs, same outputs as the worked example above. If a number on this site disagrees with the library, that disagreement is the bug and it is the highest-priority mail the contact page receives. The formula sheet lists every formula with each symbol named; the editorial policy covers how corrections are handled.
Where the methods come from
None of the statistics here is original. Removing a margin under an insider-trading model is Shin’s.[1][2] The favorite–longshot bias the power method corrects for is documented across decades of market data.[3][4] The size-weighted order-book estimator is the micro-price.[5] Merging forecasts in log-odds is a logarithmic opinion pool.[6] Scoring a venue’s calibration uses the Brier score.[7]
- Shin, H. S. (1992). “Prices of State Contingent Claims with Insider Traders, and the Favourite-Longshot Bias.” The Economic Journal 102(411), 426–435.
- Shin, H. S. (1993). “Measuring the Incidence of Insider Trading in a Market for State-Contingent Claims.” The Economic Journal 103(420), 1141–1153. – the Shin de-vig implemented in step 2.
- Thaler, R. H. & Ziemba, W. T. (1988). “Anomalies: Parimutuel Betting Markets: Racetracks and Lotteries.” Journal of Economic Perspectives 2(2), 161–174. – the favorite–longshot bias.
- Snowberg, E. & Wolfers, J. (2010). “Explaining the Favorite–Longshot Bias: Is It Risk-Love or Misperceptions?” Journal of Political Economy 118(4), 723–746.
- Stoikov, S. (2018). “The Micro-Price: A High-Frequency Estimator of Future Prices.” Quantitative Finance 18(12), 1959–1966. – the size-weighted book estimator of step 1.
- Genest, C. & Zidek, J. V. (1986). “Combining Probability Distributions: A Critique and an Annotated Bibliography.” Statistical Science 1(1), 114–135. – the logarithmic opinion pool of step 4.
- Brier, G. W. (1950). “Verification of Forecasts Expressed in Terms of Probability.” Monthly Weather Review 78(1), 1–3. – the calibration score behind step 5’s measured weights.
Frequently asked questions
How does Teacher's Bet calculate a fair line?
One price is taken from each qualifying venue, the venue’s margin is removed by de-vigging, the resulting probabilities are pooled in log-odds space, and the pooled value is converted back to a probability and an American price. Every venue is then graded against a benchmark computed with that venue excluded.
Which de-vig method does Teacher's Bet use?
The displayed fair line runs all three de-vig methods on every market, proportional, power and Shin, and keeps the most generous grade, printing the method that produced each number beside it. That max-of-three selection is the default; any single method is selectable, on the board as well as in the calculators. Grading is stricter on purpose: closing-line grades and the venue calibration ledger are computed with the proportional method alone, the one a reader can reproduce by hand, so the archived record keeps a single consistent ruler. The spread between the three methods on the same line is kept as a data-quality signal: a line the three methods disagree about earns less weight.
Why use the microprice instead of the mid on an order book?
The mid ignores how much size rests on each side. The microprice weights each quote by the size resting on the opposite side, so it leans toward the thinner half of the book, the side about to be taken. On a book with 400 contracts bid at 62¢ and 1,200 offered at 63¢, the mid is 62.5¢ and the microprice is 62.25¢.
When does Teacher's Bet refuse to publish a fair line?
A market with fewer than two qualifying venues is refused on the published board: the panel prints the reason, a venue count and a label, instead of a fair line, and an analysis failure degrades to the same refusal rather than an error. A market with no usable quote is left blank, never interpolated or carried forward.
Where venues disagree beyond the wide-dispersion threshold, the disagreement is published alongside the number rather than smoothed away, and confidence is capped at 0.45 when no anchor-tier venue is in the pool.
What is the suspect band?
Any quote computing to 15¢ or more of expected value per dollar staked is treated as a stale or erroneous line rather than an edge, and is left unflagged. The constant is EV_SUSPECT_PER_DOLLAR = 0.15. Real edges in a liquid two-way market are worth a few cents per dollar, not fifteen.
Is the fair value the same as the break-even price?
No, and they are reported as separate numbers. Fair value is the pooled probability with no margin in it. The break-even is the probability a position needs in order to cover the price plus the fee. For a Kalshi taker buying at price P, the break-even is P + 0.07 × P × (1 − P), so a 50¢ contract breaks even at 51.75%.
Does the board rank by anything besides EV?
The EV lane sorts by the pooled EV figure. A flag at a venue you have ticked sits above one you have not. A second shelf sorts a row below the rest when its side’s pooled fair line fell 0.5 points or more in the last hour (three hours for a game more than 14 days out), measured with the proportional de-vig and the flagged venue left out. Neither shelf changes an EV, a flag or an alert.