An inquiry into sustained competitive performanceNew paper · v0.9.5 · September 2026

The streak grows.The edge narrows.

A winning record can conceal a growing performance shortfall. Across college basketball and soccer, deeper winning streaks are associated with weaker performance under the tested benchmarks.

Research by James Castranova · Open-access preprint
Measures performance in games entered on a streak, including the game that ends it.

COLLEGE BASKETBALLBEYOND THE BENCHMARK
The discovery pattern by entering winning-streak depthAdjusted residual excess in points: 1 to 2 wins, plus 0.2749. 3 to 5 wins, minus 0.1249. 6 to 8 wins, minus 0.6421. 9 or more wins, minus 1.7444. These are band contrasts beyond the rating-based memoryless reference, not raw margins. benchmark +0.27−0.12−0.64 −1.74 1–23–56–89+
Consecutive wins carried into the next game. Points above or below the specified memoryless reference. See the full construction.
44,091 gamesCollege basketball discovery across eight seasons
14 football readoutsAdverse goal-differential direction in every readout
Open to scrutinyFull paper, claim status, and unresolved questions
All glory is fleeting.Patton (1970)

Everyone knows a streak eventually ends.
What changes as it gets there?

The hot-hand question asks whether success predicts more success. This research asks what happens to competitive performance as consecutive wins accumulate. A team can keep winning while the performance behind its record becomes less dominant.

Some apparent decline comes from how streaks are selected and counted. The research measures that baseline first, then asks what remains beyond it.

A better record.
A widening shortfall.

At shallow depth, adjusted performance is above its reference. At nine or more entering wins, it is substantially below.

Performance beyond the memoryless reference

Adjusted band excess relative to depth zero, in points. Discovery data.

Adjusted residual excess

The decline built into this no-duration-memory construction has already been subtracted. These are adjusted comparisons across entering-depth bands, including terminating losses. They are not raw average margins or a path traced inside every individual streak.

Exact values and simulation counts
Corrected band authority, paper §5.3
Entering winsObservedFloorExcessTail / 500
1–2−0.0328−0.3077+0.27491
3–5−0.8895−0.7646−0.124992
6–8−1.8769−1.2348−0.64212
9+−3.6314−1.8870−1.74440

Counts are reported simulation tails. Zero of 500 is not zero probability. Rounded columns can differ at the last digit.

Paper, pages 9–10
−1.74points of adjusted residual excess at nine or more entering wins

The baseline explains a lot.
It doesn’t explain all of it.

Across eight discovery seasons, the pooled KenPom residual falls about 0.27 points per additional entering win. Rebuilding the streaks under a model with no duration memory reproduces about 0.18 of that decline.

About 0.095 points per win remain beyond that benchmark. None of 500 simulated replicates was as adverse. The loss-streak comparison moves in the opposite direction and also exceeds its own floor.

Read the statistical specification

88,182 team-game rows, 44,091 games, and 2,851 team-seasons. The pooled flagship uses team-season fixed effects and dyadic cluster-robust variance. KenPom depth coefficient −0.27426, SE 0.02089, t −13.13. The 500-draw floor is −0.17933 and excess −0.09493.

The band analysis is a separate estimator using depth-band indicators, a depth-zero reference, site indicators, and three absorbed fixed-effect sets. It should not be read as the pooled coefficient multiplied by depth.

All completed games define the CBB clock before analytical eligibility is applied. Non-Division-I outcomes are held as declared in the simulation. The benchmark remains a specified model, not every possible regression or selection account.

Paper, §§4–5

The scoreboard matters

At 9+, the rating-reference result separates into −0.742 points of realized margin and +1.003 points of expected margin. Their difference is the −1.744 residual excess.

A reserved season agreed in direction

The 2025–26 win-side excess was −0.0788. The frozen directional criterion was met, with weaker simulation tails than discovery. The full depth pattern was not evaluable.

One opponent. Two histories.

Both competitors bring their own form and streak history. The opponent-state adjustment leaves the focal depth association largely intact. It does not measure a latent burden on either side.

Different leagues.
The same adverse direction.

Soccer now stands beside basketball as a principal finding. The original Premier League result has expanded across leagues, measurement systems, and time periods.

1414readouts

Goal-differential coefficients point adversely in every governed football readout.

Eight football environments. Thirteen unique league-period cells plus one overlapping EPL measurement-system bridge. This is directional recurrence, not 14 independent confirmations.

Read the soccer findings

Where the football pattern appears

Number of Understat league-period cells with an adverse coefficient. Ten cells across five leagues and two windows.

MEASUREADVERSE DIRECTION / 10
Goal differential10/10
Goals scored by the team10/10
Goals scored by its opponent9/10
Team finishing residual10/10
Expected-goal differential9/10
Team expected goals9/10
Opponent expected goals9/10
Opponent finishing residual7/10

Premier League / La Liga / Bundesliga / Ligue 1 / Serie A

Adverse means lower team scoring or higher opponent scoring, as appropriate. Finishing residual is goals minus expected goals. These counts describe signs, not significance in every cell.

The higher-level result persists. The detail changes.

Across the two EPL windows, the team’s scoring coefficient stays close: −0.043 and −0.050. The expected-goals and finishing-residual contributions change substantially. The same scoring direction can register through a different statistical decomposition.

A proposed common explanation did not confirm.

Opponent finishing residual looked like a shared feature in discovery. Fresh seasons did not confirm it as the conserved carrier. Goal differential remains recurrent. That particular explanation does not.

The original EPL test, the league clock, and the limits

The locally replicated Football-Data EPL anchor used a frozen entering-depth instrument. Its coefficients were −0.176187 for shot differential, −0.099589 for shots-on-target differential, and −0.077780 for goal differential. Corresponding fractions at least as adverse were 0.006, 0.011, and 0.000.

The ordered state is consecutive league wins. A league draw or loss resets it. The target match is measured whether it ends in a win, draw, or loss. Team-season fixed effects, pre-match opponent season-to-date goal-difference strength, venue, rest, and a within-team-season permutation reference form the declared test.

Understat EPL Window A covers 2014–15 through 2021–22, with 2,251 observations. It overlaps the earlier source and is a measurement bridge. Window B covers 2022–23 through 2024–25, with 853 observations. Continental leagues have discovery and confirmation windows.

The full football record includes the EPL and England’s three lower divisions plus the four continental leagues. Sparse deep cells prevent a general claim that soccer reproduces basketball’s four-band depth pattern. Cup results and all-competition streaks are different objects from this declared league-only clock.

Paper, §9.1 and §9.3

Success is an outcome.
Sustaining it is the question.

Entropressionnoun · the observed form

RDT proposes that sustained ordered competitive performance compresses as ordered duration accumulates. Statistical channels register that compression as measurable deformation.

In plain English: keeping a successful competitive state going becomes more difficult as it ages. The proposed burden can register differently in different sports.

Entropression names the adverse deformation measured beyond the appropriate benchmark. It does not identify the cause.

James Castranova’s interpretation is that this recurring form reflects entropic pressure on sustained competitive order. The paper measures performance, not entropy. Whether the form warrants a natural-law claim remains a question for evidence and independent scrutiny.

The contest belongs to both sides

Each competitor brings its own history and capacity. A single team’s statistics cannot stand in for the whole duel.

Skill may sustain the performance

The theory allows a capable team to compensate as difficulty grows. The present findings do not measure compensating effort or prove that explanation.

The form need not look identical

Basketball, soccer, and tennis have different clocks and measurements. A shared claim must survive tests appropriate to each, with failures recorded.

The performance changes.
Everything doesn’t get worse.

The basketball findings resist a simple story about sloppy play. They make the question more specific.

The open accounting question: at 9+, the rating-based reference gives a −0.742 realized-margin excess. A separate box-score reference gives +0.080. The references have not been reconciled, so the scoring components cannot yet be presented as the explanation for that −0.742 shortfall.

01 / THE MIXED RESPONSE

Nine of ten selected measures point favorably.

Across all four depth bands, nine of the ten previously selected process channels move in ordinarily favorable directions for the streaking team. Lower turnovers and better foul control are among them. These are descriptive point estimates on common support.

02 / THE BILATERAL EXCEPTION

Credited assist rates fall on both sides.

The streaking team’s assist rate is the exception to that favorable pattern. Its opponent’s assist rate also falls relative to its reference. A credited-assist result alone does not establish less passing, weaker teamwork, or a cause.

03 / THE SCORING IDENTITY

Two-point scoring weakens. Other scoring offsets it.

Under the exploratory box-score reference at 9+, the two-point differential contributes −1.108 points, three-pointers +0.659, and free throws +0.529. The net is +0.080. That accounting is exact within this construction.

Read the channel analysis

A finding should
take questions.

The claim is strongest when the benchmark, the alternative explanations, and the unanswered questions are visible.

Tennis · A third positive sport

Later wins take more sets.

Across 2,825 eligible streaks, a frozen lifecycle test found that later-phase wins required more sets after the declared structural adjustment. None of 2,000 null replicates reached the observed residual slope.

This uses early, middle, and late thirds of completed streaks of at least six wins. It is not basketball’s entering-depth test. Tennis’s qualifying, walkover, retirement, and missing-record audit remains open. Paper §9.2

Isn’t this just regression to the mean?

Regression and streak selection explain a substantial part of the raw basketball decline. The specified memoryless construction reproduces roughly two-thirds of the pooled KenPom coefficient. A remaining adverse excess is the finding.

Adding accumulated overperformance strengthens the depth coefficient on KenPom and BPI. Torvik behaves differently: its joint depth coefficient turns slightly positive. These results constrain specific alternatives. They do not rule out every possible regression or selection model.

Paper §§5.2, 5.5–5.6
What about fatigue, recent form, or the schedule?

Season-week controls leave the depth coefficient essentially unchanged. Days of rest and games in the preceding nine days do not materially absorb it. Those are measured proxies, not an exhaustive test of fatigue.

Recent three-game form reduces the raw depth coefficient by about 12% under KenPom and 20% under BPI. A large association remains. That is attenuation in a model, not a causal allocation of 20% to recent form and 80% to Entropression.

The specified opponent-state model also leaves the focal depth association largely intact. Unmeasured roster changes, quality drift, and opponent adaptation remain possible contributors.

Paper §6
Does this happen inside every individual streak?

That has not been established. The main basketball result compares games within the same team-season. It does not yet isolate deterioration within individual streak episodes after excluding differences in which episodes reach each depth.

A team-season fixed effect removes differences between team-seasons. It does not remove every difference between streak episodes within that season. The confirmatory within-episode question remains open.

Paper §11.1
Did the result hold up on reserved data?

In 2025–26, the frozen pooled directional criterion was met. The win-side excess was −0.07876, with 48 of 500 replicates as adverse. The loss-side excess was +0.07216, with 71 of 500 as extreme. Ten selected channel directions were concordant, without a claimed independent-channel joint probability.

The four-band lifecycle prediction was not evaluable. The 9+ band had 213 rows and 46 clusters, below the frozen floors of 500 and 50. The 6–8 band also lacked enough rows.

The specification was frozen before analytical opening of an already-played season. That does not mean nobody had watched or discussed those teams. The 2026–27 specification was frozen before the new season.

Paper §8
Has soccer reached basketball’s level?

Soccer has completed the development sequence called for in v0.9.4 and is now a principal empirical pillar. It has a prospective EPL anchor, a second measurement system, later-period evidence, and a completed continental atlas.

The strengths differ. Basketball supplies the richer deep-band shape and rival-explanation battery. Soccer supplies broader recurrence across football environments. Sparse deep soccer cells do not support a claim that it reproduces basketball’s full depth pattern.

Paper §§9.1–9.4
Does this establish a natural law?

The evidence supports a recurring duration-indexed performance form under the tested constructions. A universal law has not been established, and no physical or information-theoretic entropy is measured.

An empirical law need not wait for a complete causal mechanism. It does need a precise scope, tests capable of contradicting it, and evidence that the apparent regularity is not an artifact of the construction. Independent reproduction and the unresolved identification questions matter here.

The program’s non-confirmations and powered nulls belong in that assessment. They cannot become positive results merely because different sports express performance differently.

Paper §§10–12
Can this tell me when a team will lose?

The present evidence does not establish a universal breaking point, a game-by-game forecast, or a demonstrated betting advantage. It describes performance across duration states. Any predictive product would need its own prospective validation.

The same distinction applies to overconfidence. It is a possible contributor, but the current record does not identify it as the cause. Labeling a result “pressure” or “confidence” does not measure the mechanism.

The complete sports record

Evidence standing across the seven sports in paper v0.9.5
SportRecorded standingWhat that means
College basketballDiscovery + supported holdout concordanceSigned pooled holdout criterion met. Full lifecycle holdout not evaluable.
SoccerGrade A EPL anchor + completed atlasFourteen adverse goal-differential readouts. Distinct clocks and limited deep-band coverage.
TennisGrade A lifecycle resultConfirmed within its score-thirds instrument. Construct audit remains open.
NHLGrade B · not confirmedThe real holdout did not clear its frozen gate.
NFLGrade CA pre-holdout feasibility gate failed. Reserved data remain unopened.
College footballGrade CDiscovery selection and pre-holdout feasibility limit the finding.
MLBHistorical Grade D + newer powered nullThe well-powered starting-pitcher null is reported directly, without assigning it a new grade.

Grades describe frozen test architecture and disposition. They are not a universal-law score. Full record, paper §9.4. Repairs, retractions, and instrument history appear in the paper’s technical notes.

Entropy
to the Mean.

Winning Streaks Carry a Rising, Measurable Cost: Evidence from College Basketball and Six Additional Competitive Substrates.

VERSION 0.9.5SEPTEMBER 8, 202632 PAGESOPEN ACCESS

Preprint, not yet peer reviewed. Free to read and share with attribution under CC BY 4.0.

Regression Duel Theory

Entropy to the Mean

Winning Streaks Carry a Rising,
Measurable Cost

James Castranova
Independent researcher

A research program on sustained ordered competitive performance, its changing form, and the evidence that might explain it.

10.5281/zenodo.22664635
Citation, provenance, and reproducibility

Castranova, J. (2026). Entropy to the Mean: Winning Streaks Carry a Rising, Measurable Cost: Evidence from College Basketball and Six Additional Competitive Substrates. Preprint, version 0.9.5. https://doi.org/10.5281/zenodo.22664635

BibTeX
@misc{castranova2026entropy,
  author = {Castranova, James},
  title = {Entropy to the Mean: Winning Streaks Carry a Rising,
    Measurable Cost: Evidence from College Basketball and
    Six Additional Competitive Substrates},
  year = {2026},
  month = sep,
  note = {Preprint, version 0.9.5. Not yet peer reviewed.},
  doi = {10.5281/zenodo.22664635},
  url = {https://doi.org/10.5281/zenodo.22664635}
}

This site transcribes the reported results in v0.9.5. It does not represent a new execution or independent replication of the underlying analyses. Source artifacts, their identities, and execution scope are documented in paper §14.

The paper describes a release bundle containing the compliant band result, selected-channel result, frame manifest, and joint scoring replicate arrays. The full source datasets and full 36-channel basketball output do not accompany the review manuscript. Consult the DOI record for available release files.

PDF SHA-256: baa85d922364d3a420c5a1aab487de684bd7c047bbffdf8e4ab311a5e47efaa3

Provenance and availability, pages 26–28

Questions about the construction, alternative explanations, or independent replication are welcome.

Contact James Castranova