A Number the Market Cannot Produce
Every few months another platform launches with "live open interest" on the feature list, and every time I see it I want to ask the same question: live from where? There is no feed to be live from. Open interest does not stream. It has no order book, no tape, no tick. OI is an accounting figure, and like any accounting figure it only becomes accurate after the books close and someone reconciles them.
Here is the mechanical gap. When two parties trade an option, the print hits OPRA (the consolidated options tape) within milliseconds: price, size, exchange, timestamp. What the print does not carry is intent. Did that trade open a new position or close an old one? That single missing bit determines whether OI goes up, goes down, or stays flat. The OPRA feed simply does not contain it. Working it out requires knowing the prior positions of both the buyer and the seller, and positions are only truly netted in the OCC's overnight clearing cycle. The closest thing to an exception: orders carry a self-declared open/close marker, and exchanges sell aggregated summaries of it. That is Cboe's Open-Close data, including intraday versions at 10-minute and 1-minute snapshots. A tier deeper sit the trade-by-trade execution records those summaries are built from. It is the best labeled data that exists, and serious models train on it. But it is self-reported rather than netted, bucketed rather than account-level, per-exchange rather than consolidated, and paid. A rich input for an estimate, not a live count.
So a true live count of outstanding contracts, built from public data, can not be done. And I mean can not, in the structural sense. This has nothing to do with engineering skill or budget. The required information does not exist in any public source during market hours.
The four possible combinations for any options transaction:
| Buyer intent | Seller intent | OI change |
|---|---|---|
| Buy to Open (BTO) | Sell to Open (STO) | +Volume |
| Buy to Close (BTC) | Sell to Close (STC) | −Volume |
| Buy to Open (BTO) | Sell to Close (STC) | 0 (no change) |
| Buy to Close (BTC) | Sell to Open (STO) | 0 (no change) |
OI only moves when both sides open or both sides close. Mixed trades (one side opening, one closing) leave OI untouched, and in liquid markets they are the most common case.
The tape can not tell the four rows apart.
What the OCC Does Overnight, and Why There Is No Shortcut
The number that shows up in your chain the next morning comes out of a multi-stage reconciliation run by the OCC after the close. Every clearing member firm submits its end-of-day position records. The OCC cross-references those submissions against each other, chases down discrepancies, catches settlement failures, and produces a settled OI figure for every contract. That figure has to be legally auditable, because the next day's margin calculations sit on top of it.
The process takes hours because reconciling records across dozens of firms with different internal systems is genuinely slow work. Simple aggregation would be fast. Reconciliation isn't. And the thoroughness is exactly what makes the morning number trustworthy.
The same logic runs in reverse: the reason the overnight figure is authoritative is the reason nothing equivalent can exist during the session. You would need the account-level position data, and that data only comes together after the close.
Exchange-reported OI carries the full weight of that reconciliation. An "updated" OI figure shown at 1 PM carries none of it. When a platform renders both in the same font, the same chart format, the same apparent precision, it is dressing an estimate up as a measurement. Most platforms do exactly this, and most users never learn the difference.
GEX = OI × Gamma × Contract Size × Spot² × 0.01
Whatever error lives in the OI input flows straight through, amplified by gamma near the money. A 10% OI error at an ATM strike is a 10% error in the GEX dollar figure there. On something like the JPM Collar (40,000+ contracts at a single strike), that is hundreds of millions of dollars of phantom or missing hedging flow.
So What Are the "Live OI" Platforms Actually Doing?
To be fair to the vendors: they are not making numbers up. They are running an inference chain, and the chain is always some version of the same four steps.
Start with last night's OCC-reconciled OI. That is the verified anchor, and everything after is an adjustment layered on top of it. Then watch the OPRA tape, which is the only options data that genuinely exists in real time. For each print, classify it: opening or closing? The standard starting point is bid/ask position. A trade above the midpoint reads as buyer-initiated and gets tagged as probably opening; below the midpoint, probably closing. Finally, accumulate. Sum the inferred opening volume minus inferred closing volume across the session and add the total to the overnight baseline.
The output looks like a live OI number. Since any single trade of 500 contracts could move true OI anywhere from −500 to +500, the classifier works in probabilities, not certainties.
What it actually is: yesterday's verified figure plus a running guess about net new positioning, correct only to the degree that step three guessed intent correctly.
Where P(Open) + P(Close) + P(Mixed) = 1
P(Open) adds +Volume contribution
P(Close) adds −Volume contribution
P(Mixed) adds 0 contribution
+ Σ [Volume_i × (P(Open)_i − P(Close)_i)]
for all trades i up to time t
The whole modeling problem collapses into one question: how well can you estimate P(Open) and P(Close) for each individual trade, using only what is observable on the tape?
Three Ways to Guess Intent
Classification is where all the intellectual effort goes, and where all the error is born. Three approaches exist, roughly in order of sophistication and cost.
A. Lee-Ready / tick rule
The academic workhorse. Classify the aggressor from price: above the midpoint means buyer-initiated, below means seller-initiated, and map aggression to opening or closing intent. Cheap and fast. In equities it works reasonably well on average. In options it gets noisy, because SPX options constantly trade at mid or cross the spread inside multi-leg structures, and a big print at the ask might be a dealer closing a hedge rather than anyone going long.
B. Signed volume plus Greeks heuristics
Layer in moneyness, DTE, time of day, and IV behavior. A large OTM call bought alongside an immediate IV pop is probably opening. A near-expiry deep ITM option being sold is probably closing or rolling. OPRA strategy codes flag multi-leg trades (spreads, condors, rolls) so they can be decoded separately. More accurate than the tick rule, but now you need a tick-by-tick IV surface.
C. Delta-hedge flow reversal
Watch the futures. If a large options print is followed within roughly ±500ms by a correlated ES trade in the opposite direction, that is a dealer hedging a customer trade, which strongly suggests the customer opened a position. Very accurate for institutional blocks. Also very expensive: co-located futures tick data and sub-second timestamp alignment across two markets.
Production systems run all three in parallel and weight the outputs by trade type: tick rule for small retail prints, heuristics for mid-size flow, delta-hedge reversal for blocks (where the futures confirmation signal is strongest). Even combined, the output is still a probability, never a count.
There is also a route that skips inference entirely: buy the open/close markers from the exchange. Cboe sells its Open-Close summaries end-of-day and at 10-minute or 1-minute intraday intervals, and the SPX positioning platforms charging north of $300 a month are built on exactly this feed. One level below the summaries sits the Enhanced Trade-by-Trade Execution Detail dataset: the individual execution records themselves. Each print carries its buy/sell and open/close markers, participant capacity, and the NBBO at the moment of execution (C1 only so far, delivered T+1). That is not a dashboard input; it is the dataset you would use to train and validate a classification model against ground truth. For SPX (which trades only on Cboe) that is a genuinely complete view of marked flow: a real data-quality edge over anything inferred from the tape. Better inputs are still not the same thing as a better forecast (more on that at the end of this page). The asterisks: the markers are self-reported at order entry, the data arrives in aggregated buckets a few minutes after each interval closes, and for multi-listed names Cboe's exchanges see only a quarter to a third of total volume. The model write-up maps the competition in detail, prices included.
What the Full Stack Costs
Standard OHLCV data will not get you anywhere near this. A viable intraday OI estimator needs several expensive, high-bandwidth feeds, plus historical clearing data to train against.
Essential
| Data | Why needed |
|---|---|
| OPRA full feed | Every print: size, price, exchange, condition codes, strategy flags |
| L2 options quote feed | Bid/ask at the exact millisecond of each trade, to classify the aggressor |
| ES/SPX futures tick data | Detect delta-hedging flows that confirm customer opening trades |
| Prior day OI by strike/expiry | The baseline anchor everything else adjusts |
| IV surface tick data | IV spikes and drops help confirm opening vs closing intent |
| Historical marked open/close data | Actual opening/closing activity by strike, needed to train and validate the model (Cboe's Open-Close files, or its Trade-by-Trade Execution Detail records for per-trade ground truth) |
Very useful
| Data | Why useful |
|---|---|
| OPRA condition codes | Flag spreads, multi-leg trades, and on some venues broker-reported open/close designations |
| Block trade tape | Institutional prints dominate OI changes; modeling them separately helps |
| ETF options flow (SPY, QQQ) | Sentiment proxy that can lead SPX positioning |
The OPRA full feed alone runs $10,000 to $30,000 per month for a professional subscription. Historical OCC clearing data with open/close breakdowns is sold separately and costs thousands per year. Processing tick-level data across every SPX strike in real time needs co-location or serious cloud infrastructure. Nobody builds this over a weekend, which is worth remembering the next time a $30/month subscription service advertises live OI.
From Tick Feed to OI Estimate
Two ways to model it:
Daily-level regression (simpler)
Aggregate intraday trade features into daily buckets (share of volume at the ask, net IV change during high-volume periods, morning vs afternoon volume ratio) and train XGBoost or a Random Forest to predict end-of-day ΔOI. You do not get tick-by-tick OI, you get a daily projection that firms up as the session progresses. Easier to build, easier to validate, and honestly a better fit for GEX that refreshes every 15 minutes.
Trade-level classifier (complex)
Train on individual trades where actual open/close flags are known (CBOE LiveVol historical tick data, for example). Then deploy in real time: every print gets a P(Open) and P(Close), summed into a running estimate. Needs labeled training data and real infrastructure. It also produces error bars that (in my experience) no platform ever publishes.
The full pipeline:
ES futures ±500ms window
IV surface tick
The smoothing layer earns its keep. Classification errors accumulate over thousands of trades. So when official EOD OI publishes each day, the gap between estimate and truth gets fed back to recalibrate the classifier's priors for the next session. Skip that step and drift from true OI can hit 10 to 20% by mid-session on active days. That is enough to mislocate GEX structural levels entirely.
Where the Estimates Fall Apart
Under quiet conditions the inference chain holds up tolerably. The trouble is that the sessions where it degrades most are exactly the sessions where you would most want accurate OI. Three conditions do most of the damage.
Volume swamps the baseline. The estimate's accuracy depends on the intraday adjustment staying small relative to the overnight anchor, because every classified print adds probabilistic error. When intraday volume climbs past roughly half of overnight OI, which it does routinely on heavy 0DTE sessions and during vol expansions, the compounded error gets large enough to shift estimated OI at key strikes by amounts that matter. Pure math, no special event required.
Spreads print as separate legs. A vertical spread, a calendar, a ratio structure: each leg hits the tape individually with its own bid/ask context. The classifier scores each leg alone. But the OI impact of a spread is nothing like the sum of its legs. Classifying legs independently systematically overstates the OI change from spread activity, and institutional flow is mostly spreads.
The 0DTE closing wave. In the last thirty to sixty minutes of an expiration session, market makers flatten gamma books and traders close to dodge assignment. The opening/closing mix lurches toward closing, mechanically. A model calibrated on full-session averages will overestimate OI in precisely the window where end-of-session risk management needs it most. Watch any expiration Friday afternoon and you will see the volume signature.
Beyond those three, a set of quieter failure modes trips up even well-funded teams:
- Rolls: selling the near expiry and buying the far one nets to zero OI change but prints as two separate trades. Strategy codes help when populated, which they often are not for complex structures. Score a roll as two opens and you have invented OI that never existed.
- FLEX and ex-pit options: customized off-exchange terms do not always reach OPRA promptly, producing large unexplained OI jumps at the daily close that no tape-based model saw coming.
- Early exercise: deep ITM SPX puts get exercised early when carry cost exceeds extrinsic value. OI drops with no trade signature at all until the OCC reports it overnight.
- Internalized MM flow: some market maker trades match internally and either skip the tape or arrive delayed, forcing retroactive corrections invisible to anyone watching the live feed. And inter-dealer flow generally, one desk flattening against another, prints twice while changing end-user OI not at all.
The Poor Man's Version
Without OPRA access or tick data, you can still do better than carrying yesterday's OI unchanged. A crude proxy captures the three dominant skews in opening/closing behavior:
# on a typical day, but it varies a lot
open_ratio = base_open_rate
open_ratio += otm_skew(moneyness)
# OTM options are more likely opening
open_ratio += dte_skew(days_to_expiry)
# Very short DTE skews toward closing/rolling
open_ratio += tod_skew(time_of_day)
# First and last hour skew toward closing
estimated_OI_change = volume * (2*open_ratio - 1)
# +volume if all opening, -volume if all closing
The three skews, and why they exist:
Moneyness
OTM options get disproportionately bought to open (speculation and tail hedges). Deep ITM options lean toward closing, holders rolling out or taking delivery.
DTE
One or two days to expiry skews heavily toward closing and rolling. Seven days and beyond skews toward fresh opens as institutions build forward hedges.
Time of day
The first and last 30 to 60 minutes show elevated closing and rolling as traders flatten around the open and ahead of expiration. Mid-session flow is more balanced.
What the Estimates Are Actually Good For, and Why I Do Not Use Them Here
None of this makes intraday OI estimation worthless. Done well, it tells you which strikes are seeing genuine net new positioning versus active unwinding, flags a fresh structural level building that was not in the overnight data, and gives an early directional read on where gamma is migrating during the session. Those are real signals and I would happily take them as color.
What the estimates can not support is precision. A GEX calculation that depends on exact OI counts, a structural level pinned to a single strike, a hard number on hedging pressure at a specific price: for any of those, an estimate with an undisclosed error rate is the wrong input. Directional guide, yes. Measurement, no. And the gap between the two widens precisely when it matters, in the final hour of a 0DTE session, when position-management flow drowns out directional flow.
That asymmetry is the whole reason GEX Metrix anchors everything to verified end-of-day OI. Yes, it is hours old by the open. It is also correct. Everything the dashboard layers on top of it is an estimate, and it is labeled as one, which is the part I refuse to compromise on.
In the entire chain described on this page, the OCC's overnight figure is the only number that was measured rather than modeled.
And staleness costs less than you would think, because the positions that generate structurally meaningful GEX change slowly. As covered in the OI vs GEX article, a JPM Collar position of 40,000+ contracts does not roll intraday. The overnight OI on that position is accurate and stable, and the level it creates persists all session. Intraday estimation adds information at the margins: smaller position changes, 0DTE churn, aggressive retail flow. Real effects, much smaller notional. There are sessions where live tracking would genuinely help (Fed days and CPI prints, 0DTE expirations where OI mechanically bleeds to zero, the Monday after OPEX when new OI builds from a near-empty book). I will grant the estimators those days. For locating the Call Wall, Put Wall, and Zero Gamma level, the overnight number wins most days.
As far as I know, no other platform has a validated, publicly documented intraday OI model with a published error rate for SPX. We do, and the write-up is linked in the next section, because I am about to ask you to demand it from everyone else. So when someone advertises "live OI," skip "is it live?" (it isn't). Ask instead: what is the model's accuracy? How does that accuracy vary by moneyness, DTE, and session type? If they can not answer, you have your answer.
Where This Site Sits on the Ladder, and What Is Rolling Out
Having spent nine sections telling you what everyone else's "live OI" really is, I owe you the same transparency about mine.
The Free tier is not naive gamma. A naive dashboard takes last night's OI and shows you the same frozen map all day. GEX Metrix runs an intraday OI model on top of the verified overnight base, trained on how positioning actually migrates through a session, and you can flip between the Model view and the Naive view with one click to see the difference for yourself. I will not claim it matches a full tick-level classification engine, but it closes a meaningful part of the gap between the naive map and the approaches described above, and unlike most of the industry it is labeled as the estimate it is. It is also documented: the full technical write-up covers the model's design, training data, and error rates. Asking you to trust an unexplained estimate would make this whole page a joke.
The upgrade rolling out on the Plus tier comes straight from this page's hierarchy: a live classification model. Every trade on the tape checked against the exact bid and ask at the moment of execution, aggressor side inferred, multi-leg spreads and condition-code noise filtered out. The Lee-Ready family of methods described earlier, running on live data instead of yesterday's volume. And I will be honest about what that buys you over Free: on a quiet mid-cycle day, not much. The modeled map and the live map will usually agree. The gap opens on news days, in single names with idiosyncratic flow, and in sectors where positioning turns over fast. Those are exactly the sessions where a delayed model falls behind and a live one does not, and they are the sessions the upgrade exists for.
Which brings up the question I get asked whenever this topic comes up: why not just buy dealer position data? Cboe sells end-of-day dealer open and close positions (reported rather than inferred). A handful of platforms license data stacks like that at well north of $30,000 a month, then charge roughly $250 to $720 a month for the dashboard on top. I have run the numbers on that feed more than once, and I keep arriving at the same two problems.
The first is latency. Position data sounds like the endgame and it is not quite. The finest granularity anyone realistically gets is about a minute, and on top of that sits the exchange's own processing and dissemination lag, which nobody outside the building can pin down precisely (two minutes? three?). Meanwhile a surprise headline or one large order can rearrange the flow at a strike in seconds. So "the dealer's exact position a few minutes ago" and "a good live classification estimate right now" are much closer in practical value than the labels suggest.
The second is coverage, and this one is structural. Cboe is the only exchange that sells this positioning data. SPX options trade exclusively on Cboe, so for SPX the feed is genuinely complete. But the S&P complex is much bigger than SPX. ES options trade at CME, which publishes nothing equivalent. SPY options split across more than a dozen US exchanges, of which the Cboe feed sees only Cboe's slice. A dealer-position product covers one exchange group's view of one part of the market. Tick-level classification is the only road that goes everywhere the tape goes.
Put the two together and here is the honest scorecard. For SPX today, the platforms built on Cboe's positioning data have better raw inputs than any model that starts from yesterday's OI (mine included). But better inputs are not a crystal ball, and the more candid of those vendors say so themselves: knowing the dealer book tilts the probabilities, like a weighted coin. It does not tell you what hits the tape next, at what volume, in what volatility regime. Nobody forecasts that. Meanwhile our own strike-touch study measured a real, statistically robust edge from signed gamma levels built on modeled OI. So both roads demonstrably carry an edge, and both carry the same irreducible uncertainty about future flow. Which edge is bigger, and whether the difference justifies $300-plus a month against free, is an open empirical question. I am not aware of a single serious head-to-head study using real SPX dealer positions. And it is worth sitting with that for a second: the vendors selling these dashboards have had the actual positioning data for years, they market the edge constantly, and not one of them has published what it is worth in quantified terms. We put our modeled-OI edge on the table with numbers, sample sizes, and methodology. Until someone does the same with the dealer-position feed, claims in either direction are conviction, not measurement.
As for the roadmap here: it bets on live classification, and that part is committed. It is what the Plus tier is being built around. Whether to layer the Cboe positioning feed on top for SPX, where classification plus reported positions could plausibly end up faster than the delayed positioning snapshots on their own, is the decision I have not made. I will label each piece plainly when it goes live.
Continue Learning
Two companion pieces pick up where this one stops: why the pre-market OI snapshot stays your most accurate gamma input all session even though it is hours old, and how OI and GEX answer different questions about the same positions. If the 0DTE churn described above is your daily battlefield, 0DTE and gamma gets its own treatment too.
See What the Model View Actually Changes
Flip between the Model view and the Naive view on live SPX data, and read the full technical write-up behind the intraday OI model. The core view is free.
See GEX Data The Intraday OI Model Write-Up