Skip to content

About Mainline

The honest line

Mainline is a personalized, science-based chess training program that adapts as you play, with every claim graded the same way recommendations are graded in the app.

The thesis

No training activity has been proven to cause a measured chess rating gain. Mainline helps you train smarter on the best available evidence. It never promises you a rating.

Every chess-specific study is observational or correlational: we know what strong players do differently, but we cannot prove that copying those activities will raise your rating. Mainline is built around that fact, not around hiding it.

The vision

The training-program layer, not another tool.

Mainline is not another puzzle trainer, game analysis tool, or spaced-repetition deck. This app sits one layer up: it is the orchestration layer that decides what you should work on, with which resources, and why, then revises that plan continuously as you train and play.

Engagement and progress

Built to support consistency, not to extract attention.

Progress in Mainline means training signals: whether you are showing up, completing the planned work, keeping reviews healthy, and building skill estimates with uncertainty. Rating is noisy, and no activity here is treated as proven to cause rating gain.

The engagement layer exists because consistency is part of training. It uses forgiving reminders, capped streak cycles, and competence feedback to make practice easier to resume. It does not use ads, leaderboards, shame, unbreakable streaks, or paywalled training quality.

The boundaries

What this app deliberately is not.

Every exclusion is a design decision with a reason.

  • No LLM/AI at runtimeAI plays chess poorly and invites cost, abuse, and opacity. The app is pure deterministic algorithms. You can verify every decision.
  • No competing game platformLichess and Chess.com already do that better. The app references external platforms; it doesn't replace them.
  • No hosted copyrighted contentBooks and courses are recommended and logged, never hosted. The app points you to the right resource; it doesn't steal it.
  • No social or multiplayerNo leaderboards, no chat, no shared sessions. Social comparison harms long-term motivation, and the app's goal is personal improvement.
  • No self-report skill diagnosisDunning-Kruger is real in chess (Grade A/1). The app diagnoses you behaviorally from your games and your puzzle performance, never by asking you to rate yourself.
  • No infinite streaksUnbreakable streaks are a dark pattern. They create a loss-aversion "quit moment" when the streak breaks. The app caps streaks and forgives missed days.
  • No global leaderboardsDownward social comparison (seeing yourself ranked below strangers) harms motivation for the majority of users who are not at the top.
  • No puzzle-volume chasingCorrelation between puzzle volume and rating gap is r=−0.02. Grinding puzzles without reflection or spacing doesn't help. The app prioritizes how you practice over how much.
  • No opening memorization for beginnersBeginners lose to blunders, not opening theory. Time spent memorizing lines at <1200 is time not spent on tactics and board vision.

The evidence framework

Borrowed from the board: every claim is annotated.

Every recommendation, methodology value, and claim on this page carries a grade: a placeholder can never pose as established fact.

Grade A

Strong, replicated

Used for robust, replicated findings, e.g. the retrieval-practice effect or the spacing effect.

Grade B

Suggestive, limited

Suggestive studies with limited sample size, context, or generalizability.

Grade C

Theory / best-guess

Logical inference, placeholder, or calibration estimate. Treated as a starting point, never a proven prescription.

Grade D

Myth: avoided

Popular chess-improvement advice the app actively avoids because the evidence contradicts it.

Grade answers how strong the science is. A second axis, confidence, answers a different question: how much of your own data backs this specific call to you. The same Grade-A finding can land with low confidence, as when we know spaced repetition works but you've only imported three games. Or it can land with high confidence, well-backed by your own play. The distinction keeps a band prior from masquerading as a personalised verdict.

Insufficient

Not enough of your data yet

We don't have enough of your games or reviews to make this call. The app says so plainly instead of inventing a verdict.

Low

A band prior, not your own data yet

The recommendation rests on what players at your level tend to need, not on what we've seen from you. It will sharpen as your data accrues.

Medium

Some of your own data

Partially grounded in your games or reviews. A working hypothesis, still refining.

High

Well-backed by your own data

Drawn from enough of your own play to read as yours, not as a population average.

Where the science is now

Honest about the current state.

Active methodology · research-1.4.0

This active research release encodes the approved methodology values and copy, while retaining every documented best guess and deliberate stub as evidence-labeled data.

The release is reproducible: historic programs keep the version and rationale snapshot they were generated with. The central caveat remains unchanged: no training activity has been proven to cause a measured rating gain.

Still deliberately unresolved

  • · Training-fit feedback is not evidence of chess skill and no adherence or rating effect is claimed.

Aggregate basis

None. P9 enables controlled observational export, but no Mainline aggregate has informed this release.

Methodology release history

Every change keeps its limits and rollback path.

research-1.4.0 · 2026-07-15

Adds the sparse P8 training-fit prompt policy and restricts subjective fit to a positive tie-break inside the evidence-led daily mix.

Aggregate basis: None. P9 enables controlled observational export, but no Mainline aggregate has informed this release.

Evidence changes:

  • · Added Grade C training-fit policy values and retained the Grade A boundary separating self-report from behavior.

Limitations:

  • · Prompt timing and positive tie-break effects are unvalidated product best guesses and do not establish adherence or rating effects.

Rollback: Restore research-1.3.0 if prompts or fit ordering cause operational harm; preserve 1.4.0 artifacts for replay.

research-1.3.0 · 2026-07-12

Keeps the three-item calibration and revises weekly-focus alternative copy so optional user choice is clear without weakening its evidence caveat.

Aggregate basis: None. No Mainline observational aggregate informed this release.

Evidence changes:

  • · Clarified user-facing rationale without changing its Grade C, Tier 2 evidence.

Limitations:

  • · Alternative-choice framing has not been validated for chess training adherence.

Rollback: Restore research-1.2.0 if the revised optional-choice copy is misleading.

research-1.2.0 · 2026-07-12

Shortens new-user calibration to one three-item tactical track while preserving historic assessment behavior by methodology version.

Aggregate basis: None. No Mainline observational aggregate informed this release.

Evidence changes:

  • · No evidence grade changed; the shorter calibration is explicitly Grade C.

Limitations:

  • · Three calibration items may trade completion against measurement quality and require observational review.

Rollback: Restore research-1.1.0 for new assessments; never rescore historic assessments silently.

research-1.1.0 · 2026-07-12

Adds stable weekly focus selection, confidence-gated revision, and bounded process-goal alternatives.

Aggregate basis: None. No Mainline observational aggregate informed this release.

Evidence changes:

  • · Added Grade C weekly-focus policy values without upgrading the underlying evidence.

Limitations:

  • · Focus stability and bounded alternatives are product best guesses, not demonstrated causes of adherence or rating gain.

Rollback: Restore research-1.0.0 if weekly-focus selection is operationally unsafe; historic programs retain 1.1.0.

research-1.0.0 · 2026-07-10

First research-channel methodology release for the current Phase 1 seams.

Aggregate basis: None. No Mainline observational aggregate informed this release.

Evidence changes:

  • · Published the reviewed research synthesis with its existing grades and citations; no evidence was upgraded.

Limitations:

  • · Chess-specific activity effects remain observational or extrapolated and cannot establish rating causation.

Rollback: Restore stub-0.1.0 only as an owner-reviewed operational rollback while preserving historic program versions.

stub-0.1.0 · 2026-06-21

Pre-release placeholder configuration retained for historic programs.

Aggregate basis: None. No Mainline observational aggregate informed this release.

Evidence changes:

  • · None. This was a pre-release placeholder.

Limitations:

  • · Not an evidence-complete methodology release.

Rollback: Historic only. Do not activate without owner review.