A green status is one checkpoint. After a restart, curtailment event, network interruption or maintenance window, recovery means proving that useful work returned and the operating point stayed stable.
The event ends. Power returns. Rows of miners begin to answer. The dashboard changes from red to green, and the incident channel gets quieter.
It is tempting to call the fleet back.
But green can mean only that a device is reachable or has crossed one application's definition of online. It does not, by itself, prove that the correct build and profile loaded, tuning completed, pool work resumed, submitted work was accepted, wall power settled, temperatures stopped moving, or an earlier error stayed gone.
The dashboard can tell you the miner answered. The pool tells you whether useful work returned. Time tells you whether it stayed.
That gap is the restart recovery window. It begins with an authorized recovery action and ends when the operator's declared evidence supports a stable return. Planned reboot, maintenance, power, curtailment and network events are not interchangeable. Each starts with a different question.
There is no universal recovery time. Different fleets cross relevant operating cycles differently. A network interruption does not ask the same electrical question as re-energizing a power group. Make the clock visible and the exit criteria explicit.
TL;DR
- Online status is one checkpoint, not proof of full recovery.
- Keep firmware startup, autotune when used, and final stabilization as separate recorded stages.
- Keep a configured firmware control separate from measured wall power and the site-approved wall-power ceiling.
- Verify pool connection, submitted work and accepted work as different signals.
- Restart one canary per coherent cohort before advancing a larger wave.
- Use the Restart Recovery Card to record timestamps, evidence, stop rules and ownership.
- Decide GO, HOLD or STOP from a declared observation window. Do not inherit a universal duration or threshold.
Online is one checkpoint. Recovery is an evidence chain.
Green appears early. The recovery window stays open until the declared evidence supports the next decision.
Online is a snapshot
NIST uses a useful distinction in its engineering statistics handbook: quality is like a snapshot at the beginning, while reliability is a moving picture of operation over time. The analogy fits a restart without turning mining into a generic reliability equation.
An online tile is a snapshot. It says something happened at that moment, according to that system's definition. Recovery is the moving picture.
Online is not a universal protocol state. It may mean a responding control board, recent heartbeat, reachable address, or active pool connection. Document what the exact tool means.
Uptime answers a narrow, tool-specific question. It does not prove accepted work, wall power, stable thermals, or freedom from recurring errors. A miner can submit work before its operating point stops moving.
Green is a state change. Recovery is an evidence chain.
The restart evidence chain
Use this sequence as a set of gates:
Event classified -> power or network path restored -> device reachable -> correct build and profile confirmed -> startup complete -> autotune complete if invoked -> pool connection open -> work submitted -> work accepted -> wall and machine signals observed -> recurrence checked -> steady-state decision
Each arrow is a question. It is not proof that the earlier stage caused the later result.
1. Device reachable
Confirm the exact miner or cohort, not just a row count. Record model, control-board route, current build, configured profile or limit, network identity and event type. The live VNISH data catalog maps exact model, board, build, file and install route. It does not contain the local site's power, cooling, network or maintenance state.
For context, BITMAIN lists the exact base S21 at 200 TH/s and 3,500 W power on wall at 25 degrees C. Those figures belong to that named model and stated condition. They are not a recovery threshold for all S21-class machines, another S21 model, or an individual field unit.
2. Firmware startup
Confirm that the intended build and profile are active and that expected boards and chips are visible. Preserve first-boot warnings, missing hardware, configuration drift and time stamps. Do not erase the first error just because a later screen looks clean.
The current VNISH 1.3.5 release page treats installation, autotune and final stabilization as separate stages. That distinction matters after a restart. Startup is not autotune. An autotune completion signal, when autotune was actually invoked, is not final stabilization. None of the stages promises a result.
3. Measured wall power
Record measured wall power at a named boundary using the site's approved method. Keep that value separate from firmware-reported watts and from any configured limit. A configured firmware control is a setting. It is not proof of power at the wall and it is not the site's electrical approval.
The VNISH platform exposes requested and observed operating points, watts, measured efficiency, temperatures, fan response, board state, errors and active profile as decision signals. Its displayed values are interface examples, not a performance promise or evidence from this site.
When a power group returns, electrical limits must be approved by the responsible site specialist. A green miner tile cannot approve a circuit, PDU, PSU, connector, feed or restart sequence.
4. Pool connection and work submission
A pool connection is another gate, not the finish. It can show that a channel or session exists and jobs can move. It does not necessarily show that useful work has been accepted through the required window.
The Stratum V2 Mining Protocol separates work distribution, channel opening, client share submission, success responses and error responses. Its success message includes a count of newly accepted submits and the sum of acknowledged share difficulty. The protocol makes the operational distinction clear: connection, submission and acceptance are not the same event.
Use the exact fields your pool exposes. Record pool, endpoint, worker identity, time zone and observation window. If rejects or stales are available, preserve them under their actual labels. Do not invent a substitute when a field is unavailable.
For a deeper explanation of device-side, pool-estimated and paid or credited work, use the ROI ASIC guide Your ASIC Has Three Hashrates. This article asks a different question: after an interruption, when is the miner trusted to be back?
5. Stable operating point
Accepted work returning is necessary evidence. It is not the last evidence.
Observe wall power, reported watts, observed hashrate, accepted work, rejects or stales when available, temperatures, fan behavior, board state, errors and resets across a declared window. Look for recurrence, oscillation and drift, not only a favorable end value.
The official VNISH fleet guidance gives the sequence clearly: baseline window, first boot check, early observation and an agreed evaluation window. It says the minimum interval is agreed per device cohort before the pilot. That is the right discipline here. Select a window that covers the site's relevant operating cycles, document interruptions, and explain why it is sufficient for this decision. Do not call any duration universal.
The hidden operational cost of a false recovery
A premature green light creates cost even when nobody writes a dollar figure beside it.
Equipment can be energized before useful accepted work fully returns. Technicians may revisit miners cleared too early. Repeated resets restart observation. A fleet may look recovered in aggregate while one cohort remains unstable. An unrecorded profile change can destroy the baseline comparison.
The cost also appears as uncertainty. Without first-acceptance time, the team cannot separate recovery from normal pool variance. Without wall measurement, it cannot establish the physical load. Without error timestamps, it cannot separate an old warning from recurrence.
None of this proves those costs are common or assigns them a universal price. It identifies the work that disappears when recovery is reduced to one status light.
The recovery window is the distance between powered and trusted.
The Restart Recovery Card
Use one card per canary or coherent restart cohort. Record fields only when they exist. Mark unavailable fields as unavailable.
| Field | Record | Why it matters |
|---|---|---|
| Event identity | Event ID, type, reason, start, owner | Keeps maintenance, network, power and curtailment events distinct |
| Device or cohort | Exact assets and cohort definition | Prevents unlike machines from sharing one recovery verdict |
| Hardware route | Exact model and control-board route | Anchors the supported build path |
| Firmware state | Current build, profile, configured limit, change owner | Shows what the device was asked to run |
| Pre-event baseline | Wall power, telemetry, pool fields, thermals, errors, window | Gives recovery a real comparison point |
| Site context | Power-feed group, approved wall ceiling, cooling zone, network path | Records physical and network constraints |
| Recovery authorization | Who approved re-energizing or reconnecting | Keeps safety and operations ownership explicit |
| Device online | Timestamp and exact system definition | Records reachability without calling it recovery |
| First boot check | Build, profile, boards, chips, warnings, time | Captures early configuration and hardware state |
| Autotune stage | Invoked or not, start, finish, result label | Keeps tuning separate from startup and stabilization |
| Wall power | Instrument, boundary, timestamp, value or range | Records physical demand at the named point |
| Firmware watts | Exact interface label, timestamp, value or range | Preserves the device estimate without relabeling it |
| Pool connection | Pool, endpoint, worker, first connection time | Shows the work path is open |
| Submitted work | Exact field and first observed time | Shows results are leaving the miner |
| Accepted work | Exact pool field, first acceptance, window total or series | Shows recognized work returned |
| Rejects and stales | Exact pool fields when available | Adds delivery quality context |
| Operating signals | Requested and observed hashrate, temperatures, fans, boards | Shows how the operating point develops |
| Recurrence | Errors, resets, missing hardware, timestamps | Tests whether the incident condition returns |
| Observation window | Start, end, timezone, cycles, interruptions | Makes the decision reproducible |
| Stop rules | Site-specific electrical, thermal, hardware, pool and stability rules | Defines when evidence overrides momentum |
| Decision | GO, HOLD or STOP, evidence and owner | Creates an auditable rollout gate |
The card needs a consistent record across the tools that own each signal.
An eight-step staged restart
1. Classify the event
Name what changed. A network interruption may start with path and endpoint checks. A power event requires site-authorized electrical checks. Maintenance requires confirmation of the exact work performed. Curtailment return needs the operating plan and site event history preserved. For deeper curtailment context, read Bitcoin Mining Is Flexible, But Not at Any Price.
2. Freeze the baseline and stop rules
Preserve the last valid pre-event window when available. Declare site-specific stop rules, responsible people and measurement boundaries before restarting the pilot. Missing baseline evidence is not permission to invent it. Record the gap.
3. Form coherent restart cohorts
Group machines by the variables that can change this recovery decision, such as exact hardware route, build, power-feed group, cooling zone, maintenance state and observed condition. The VNISH Global cohort guide explains why a fleet average can hide the response tail.
4. Start one canary per cohort
Restart one representative canary before the wave. Keep an untouched peer or control group where the event and site design make that possible. Do not use a miner with unresolved faults as the only representative of a healthy cohort.
5. Confirm first boot before tuning
Verify identity, expected build, configured profile, visible boards and chips, warnings and recovery readiness. If autotune is part of the supported plan, record it as a separate stage. Do not mix its start or completion with the startup timestamp.
6. Rebuild the work path
Record pool connection, job flow when visible, submitted work, successful acceptance and error fields. Align the time zone and window across miner, pool and site logs. Do not treat first acceptance as proof of steady state.
7. Observe the operating point
Compare the canary with its baseline and untouched peer through the declared window. Keep the site-approved wall-power ceiling separate from the configured firmware setting. The VNISH Ninja watt-budget guide provides the measurement discipline for that boundary.
Watch the series, not one screenshot: wall power, reported watts, accepted work, hashrate, thermals, fan response, board state, errors and resets. Record site or network interruptions rather than smoothing them away.
8. Decide before the next wave
Choose GO, HOLD or STOP for that cohort. If GO, advance a controlled batch and repeat the gates. A canary clears a next wave, not the entire fleet forever. If the evidence splits, form a separate cohort rather than averaging away the difference.
GO, HOLD or STOP
| Decision | Use when | Next action | It does not mean |
|---|---|---|---|
| GO | The canary met prewritten site criteria across the declared recovery window | Advance only its coherent cohort to the next controlled wave | Every machine, site or future restart is cleared |
| HOLD | Evidence is incomplete, windows do not align, an interruption invalidated comparison, or behavior has not settled | Preserve the current state, resolve the evidence gap and restart observation if appropriate | The firmware or hardware failed |
| STOP | A prewritten electrical, thermal, hardware, pool-work, error, reset or stability condition was reached | Contain the wave, preserve evidence and follow qualified support or recovery procedure | One universal threshold applies elsewhere or recovery is guaranteed |
Persistent reset, board, thermal or accepted-work issues may move a machine out of the recovery workflow and into triage. The ROI ASIC tune, repair or replace decision guide helps structure that next question without assuming firmware caused the issue.
What this does not prove
A completed card does not prove that firmware caused every observed change. Power quality, cooling, machine condition, pool behavior, network path, maintenance and instrumentation can also move the result.
One recovery window does not forecast another site, season, event type, pool, build or hardware age. A canary reduces uncertainty for a defined cohort. It does not eliminate uncertainty.
An accepted share does not guarantee a future payment, stable hashrate, stable power or stable temperature. A stable window does not guarantee hardware life, revenue, efficiency, recovery success or freedom from another incident.
Aftermarket firmware and changed operating profiles alter machine operating conditions. They can affect power draw, temperature, stability, component stress, hardware life and manufacturer warranty or support conditions. Results depend on the exact model, control board, silicon, PSU, power input, cooling, network, pool, ambient conditions, build and settings. Qualified personnel must approve electrical work, limits and recovery procedures.
ROI ASIC is part of the VNISH ecosystem. Use the exact supported firmware route, preserve a recovery path appropriate to the board and site, and start with one measured canary. No performance, savings, uptime, recovery, revenue or financial result is guaranteed.
The next restart starts now
Do not wait for the next incident to invent the card.
Take one recent restart. Reconstruct the event type, online timestamp, first accepted work, wall measurement, error history and steady-state window. The missing fields will show what the next recovery plan needs.
Then choose one coherent cohort and one canary. Write the stop rules before the restart. Keep the evidence chain visible until GO, HOLD or STOP has an owner.
Your miner is not back because a tile changed color. It is back when the declared evidence says the operating point returned, useful work stayed accepted, and the next rollout decision is defensible.
Review six source-linked performance reports before converting a reported TH/s, watt or J/TH figure into an economic assumption. Open the independent field reports.
Sources and data freshness
- VNISH GLOBAL, VNISH 1.3.5 release page, current page checked 23 August 2026. Used for route identity and the distinction among installation, autotune and final stabilization. No release-date claim imported.
- VNISH GLOBAL, operational platform, fleet rollout guidance and data catalog, checked 23 August 2026. Demo values are examples, not performance evidence.
- Stratum V2, Mining Protocol specification, checked 23 August 2026. Used for the separation of connection, share submission, success and errors.
- NIST/SEMATECH, Quality versus reliability, checked 23 August 2026. Used only for the snapshot-versus-time distinction.
- BITMAIN, ANTMINER S21 specification, updated 8 April 2024 and checked 23 August 2026. Figures apply only to the exact base S21 and stated conditions.
Published by the ROI ASIC Analytics Desk. ROI ASIC is part of the VNISH ecosystem. This field guide does not promise a performance, uptime, recovery or financial result.