TLDR: The easiest way to track MTG deck changes is to preserve a dated baseline, record every card added and removed, and give each revision one clear hypothesis. Compare measurable deck properties immediately, then test the relevant gameplay situations. Opening hands and goldfishing can expose mana or sequencing problems; actual games provide matchup context. Neither a win nor a loss proves that one swap worked.
If you want to track MTG deck changes usefully, do not build a diary that says “went 3–1, deck good.” Build a version log that connects an observed problem to a specific edit and a defined test. Magic contains too much draw variance, opponent variance, and decision complexity for a handful of match results to identify one card as the cause.
The goal is controlled iteration, not laboratory perfection. Preserve what you started with, change as little as practical, inspect what changed structurally, and collect the kind of evidence that matches the problem. That gives you a reason to keep, revert, or continue testing an edit instead of trusting whichever game you remember most vividly.
Start with a decklist baseline
Before making an edit, create a dated copy of the current list. You can export it, duplicate it in whatever deck-storage system you use, or save the list as a plain-text file. If you are starting from scattered notes, use the MTG Deck Builder to assemble a clean list, but keep the version log in a separate document unless you have confirmed that your chosen tool provides the revision controls you need.
Give the baseline a simple version name such as “v1.0 — 2026-09-10.” Record the format, main plan, expected environment, and recurring issue. The environment matters: a Commander deck tuned for slow creature-heavy tables is not being tested under equivalent conditions when it enters a fast combo pod.
- Version name and date
- Format and any relevant legality context
- Deck’s primary plan and expected pace
- Known problem you are trying to solve
- Full main-deck, commander, companion, or sideboard configuration as applicable
- Relevant context, such as local metagame, Commander pod expectations, or Limited event
Do not overwrite this baseline. A previous version is valuable even when you think the new build is obviously better. “Obviously” has retired many perfectly functional cards.
Create one focused change log
Each new version should answer four questions: What changed? Why did it change? What should improve? What could get worse? This turns an edit from a preference into a testable hypothesis.
For example, suppose an opening-hand review repeatedly shows that a deck cannot affect the battlefield before turn three. A useful hypothesis would be: “Replacing two expensive, situational spells with two flexible two-mana plays will reduce nonfunctional early hands without materially weakening the late game.” That is better than “trying two cool cards.”
Small revisions are easier to interpret, but one-card-at-a-time editing is not an absolute rule. Some changes only make sense as packages: adding a new color source may accompany a spell with demanding color requirements, while removing a combo piece may require replacing its support cards. Keep the revision focused on one purpose even if it involves several cards.
Copyable deck version template
- Version and date:
- Format and testing environment:
- Observed problem:
- Hypothesis:
- Cards out:
- Cards in:
- Structural changes:
- Expected benefit:
- Possible cost or new weakness:
- Test method:
- Opening-hand or goldfish observations:
- Game and matchup observations:
- Decision: keep, revert, revise, or test longer
- Reason for decision:
Separate measurable changes from conclusions
A deck edit changes objective list properties before it changes any game result. Compare those properties first. MTGApp presents its deck builder as a list-building workspace and its mana-curve analyzer as a way to inspect deck structure; the two jobs are related but distinct. The deck builder and mana-curve analyzer workflow is useful here: revise the list, then inspect what the revision actually did.
| Track immediately | Interpret after testing |
|---|---|
| Total cards and land count | Whether the deck functions under realistic game pressure |
| Mana-value distribution | Whether earlier plays improve relevant matchups |
| Colored mana requirements | Whether the mana base casts spells on schedule |
| Counts of threats, answers, draw and ramp | Whether those roles appear in the situations where they matter |
| Number of early plays | Whether they improve sequencing rather than merely filling turns |
| Sideboard cards assigned to each matchup | Whether the post-board plan trades resources effectively |
Mana curve is contextual rather than a universal formula; Wizards’ overview of mana fundamentals likewise explains curve decisions in relation to deck strategy. A lower average mana value is not automatically an improvement. Replacing a five-mana removal spell with a two-mana threat lowers the curve, but it also reduces interaction and shifts the deck toward proactive play. The useful question is not simply “Did the curve go down?” It is “Did the deck gain the early action it needed without losing an essential role?”
If the deck still feels clunky after lowering its curve, inspect other causes. Difficult color requirements, insufficient lands, weak fixing, too little cheap card selection, or a mismatch between ramp and expensive spells can produce similar symptoms. The high-mana-card diagnostic workflow helps separate those problems before every six-drop gets escorted from the premises.
Use a hypothesis-test loop
A repeatable tuning loop keeps the evidence attached to the original problem.
- Observe a recurring failure. Describe the situation, not just the result. “Missed the third land drop in several opening-hand sequences” is more useful than “the deck lost.”
- State a hypothesis. Identify what deck property you believe is causing the problem.
- Make a focused revision. Change the smallest coherent card or package set that can test the idea.
- Recheck structure. Compare land count, curve, color demands, early plays, and functional roles before playing.
- Test under reasonably consistent conditions. Use the same mulligan policy and similar matchup or pod context where practical.
- Record observations. Note whether the target problem appeared, plus any new costs created by the edit.
- Decide deliberately. Keep the revision, restore the baseline, modify the hypothesis, or collect more evidence.
There is no universal number of games or sample hands that validates an edit. A clear structural mistake may appear quickly, while a matchup-dependent sideboard choice needs broader testing. Stop when you have enough relevant observations to make the next decision—not when a magic counter reaches ten.
Match the test to the problem
Opening hands and goldfishing
Repeated opening-hand checks are useful for detecting internal consistency problems: too few castable spells, awkward color combinations, missed land drops, hands overloaded at one mana value, or sequences that rely on drawing one specific card. Label recurring failure modes and avoid drawing a conclusion from one spectacularly good or bad hand.
Goldfishing extends that inspection into early sequencing. It can show whether lands enter tapped at inconvenient times, whether ramp leads anywhere useful, and whether competing color requirements interfere with the intended line. It cannot reproduce removal, counterspells, combat pressure, hidden information, or multiplayer politics. Use the opening-hand and mulligan framework to keep decisions consistent across versions.
Actual games
Games are necessary when the question involves opponents: whether an answer is too narrow, a threat survives long enough to matter, a sideboard plan improves resource exchanges, or a Commander engine attracts more pressure than the deck can handle. Record relevant situations rather than only final results.
Win rate can be included, but treat it as context unless you have a large, reasonably comparable sample. A 2–0 evening could reflect favorable pairings, strong draws, opponent mistakes, or the changed card never appearing. Conversely, a card can solve its assigned problem during a loss. Record whether the hypothesis was tested in the game before using that result to judge the revision.
Adjust the log for your format
Commander
Record the commanders, approximate table pace, major strategies represented, and whether the edited card was drawn or tutored. Also note what happened when it mattered. Multiplayer results are especially noisy because threat assessment, alliances, seating order, and three opposing decks change the conditions. “Lost after adding more removal” tells you almost nothing; “new two-mana removal answered an early engine without consuming the entire turn” evaluates the intended job.
Constructed
Record the opposing archetype, whether you played or drew first, the game number, and the sideboard configuration. Test the post-board plan as a package. Adding three cards means little if you cannot identify what leaves the main deck or if the swap quietly damages your own primary plan.
Limited
For Limited, keep post-event notes about cards that underperformed, curve gaps, fixing, sideboard options, and picks you would reconsider. Use those observations to improve future drafts or builds. Do not assume you can freely revise a registered pool or deck during an event; the applicable event rules and procedures control what changes are permitted.
Keep personal testing separate from tournament rules
Ordinary deck iteration and sanctioned-event deck configuration are different contexts. Under the cited Magic Tournament Rules, Constructed decks generally have a minimum of 60 cards and may have a sideboard of up to 15 cards. Sideboard changes made between games must be returned to the original deck configuration before the first game of the next match. Check the current Magic Tournament Rules and event-specific instructions before changing a registered deck or sideboard.
Legality also belongs in the version log when it could affect the build. Wizards states that a card banned in a format cannot be included in a deck or sideboard for sanctioned events in that format. Record the format and date so an older snapshot is not mistaken for a currently legal list.
When to keep, revert, or test longer
- Keep the change when it addresses the stated problem, preserves essential roles, and does not create a more serious recurring weakness.
- Revert when the structural cost is already clear, the new card does not perform its assigned role, or the original version better supports the deck’s plan.
- Revise the hypothesis when testing reveals that you diagnosed the wrong cause—for example, the apparent curve problem is actually a color-source problem.
- Test longer when the relevant card or matchup rarely appeared, game contexts varied sharply, or the observations point in different directions.
The final question is not “Did this version win more tonight?” It is “What did the edit change, did that change address the original problem, and what evidence would alter my conclusion?” Preserve the baseline, log the next focused revision, and run the same loop again. Deck tuning becomes much easier when every edit leaves a receipt.