ROI ASIC Research · Restart protocol

Your Miner Is Not Back When the Dashboard Says Online

A green status is one checkpoint. Recovery means proving useful work returned and the operating point stayed stable.

A green status is one checkpoint. After a restart, curtailment event, network interruption or maintenance window, recovery means proving that useful work returned and the operating point stayed stable.

The event ends. Power returns. Rows of miners begin to answer. The dashboard changes from red to green, and the incident channel gets quieter.

It is tempting to call the fleet back.

But green can mean only that a device is reachable or has crossed one application's definition of online. It does not, by itself, prove that the correct build and profile loaded, tuning completed, pool work resumed, submitted work was accepted, wall power settled, temperatures stopped moving, or an earlier error stayed gone.

The dashboard can tell you the miner answered. The pool tells you whether useful work returned. Time tells you whether it stayed.

That gap is the restart recovery window. It begins with an authorized recovery action and ends when the operator's declared evidence supports a stable return. Planned reboot, maintenance, power, curtailment and network events are not interchangeable. Each starts with a different question.

There is no universal recovery time. Different fleets cross relevant operating cycles differently. A network interruption does not ask the same electrical question as re-energizing a power group. Make the clock visible and the exit criteria explicit.

TL;DR

  • Online status is one checkpoint, not proof of full recovery.
  • Keep firmware startup, autotune when used, and final stabilization as separate recorded stages.
  • Keep a configured firmware control separate from measured wall power and the site-approved wall-power ceiling.
  • Verify pool connection, submitted work and accepted work as different signals.
  • Restart one canary per coherent cohort before advancing a larger wave.
  • Use the Restart Recovery Card to record timestamps, evidence, stop rules and ownership.
  • Decide GO, HOLD or STOP from a declared observation window. Do not inherit a universal duration or threshold.
Restart recovery timeline

Online is one checkpoint. Recovery is an evidence chain.

01EVENTowner and type
02ONLINErecord definition
03STARTUPbuild and profile
04POOLendpoint and worker
05ACCEPTEDnamed pool field
06STABLEwall, heat, errors
07DECIDEGO, HOLD or STOP
DEVICE AND SITEPOOL AND WORK

Green appears early. The recovery window stays open until the declared evidence supports the next decision.

Online is a snapshot

NIST uses a useful distinction in its engineering statistics handbook: quality is like a snapshot at the beginning, while reliability is a moving picture of operation over time. The analogy fits a restart without turning mining into a generic reliability equation.

An online tile is a snapshot. It says something happened at that moment, according to that system's definition. Recovery is the moving picture.

Online is not a universal protocol state. It may mean a responding control board, recent heartbeat, reachable address, or active pool connection. Document what the exact tool means.

Uptime answers a narrow, tool-specific question. It does not prove accepted work, wall power, stable thermals, or freedom from recurring errors. A miner can submit work before its operating point stops moving.

Green is a state change. Recovery is an evidence chain.

The restart evidence chain

Use this sequence as a set of gates:

Event classified -> power or network path restored -> device reachable -> correct build and profile confirmed -> startup complete -> autotune complete if invoked -> pool connection open -> work submitted -> work accepted -> wall and machine signals observed -> recurrence checked -> steady-state decision

Each arrow is a question. It is not proof that the earlier stage caused the later result.

1. Device reachable

Confirm the exact miner or cohort, not just a row count. Record model, control-board route, current build, configured profile or limit, network identity and event type. The live VNISH data catalog maps exact model, board, build, file and install route. It does not contain the local site's power, cooling, network or maintenance state.

For context, BITMAIN lists the exact base S21 at 200 TH/s and 3,500 W power on wall at 25 degrees C. Those figures belong to that named model and stated condition. They are not a recovery threshold for all S21-class machines, another S21 model, or an individual field unit.

2. Firmware startup

Confirm that the intended build and profile are active and that expected boards and chips are visible. Preserve first-boot warnings, missing hardware, configuration drift and time stamps. Do not erase the first error just because a later screen looks clean.

The current VNISH 1.3.5 release page treats installation, autotune and final stabilization as separate stages. That distinction matters after a restart. Startup is not autotune. An autotune completion signal, when autotune was actually invoked, is not final stabilization. None of the stages promises a result.

3. Measured wall power

Record measured wall power at a named boundary using the site's approved method. Keep that value separate from firmware-reported watts and from any configured limit. A configured firmware control is a setting. It is not proof of power at the wall and it is not the site's electrical approval.

The VNISH platform exposes requested and observed operating points, watts, measured efficiency, temperatures, fan response, board state, errors and active profile as decision signals. Its displayed values are interface examples, not a performance promise or evidence from this site.

When a power group returns, electrical limits must be approved by the responsible site specialist. A green miner tile cannot approve a circuit, PDU, PSU, connector, feed or restart sequence.

4. Pool connection and work submission

A pool connection is another gate, not the finish. It can show that a channel or session exists and jobs can move. It does not necessarily show that useful work has been accepted through the required window.

The Stratum V2 Mining Protocol separates work distribution, channel opening, client share submission, success responses and error responses. Its success message includes a count of newly accepted submits and the sum of acknowledged share difficulty. The protocol makes the operational distinction clear: connection, submission and acceptance are not the same event.

Use the exact fields your pool exposes. Record pool, endpoint, worker identity, time zone and observation window. If rejects or stales are available, preserve them under their actual labels. Do not invent a substitute when a field is unavailable.

For a deeper explanation of device-side, pool-estimated and paid or credited work, use the ROI ASIC guide Your ASIC Has Three Hashrates. This article asks a different question: after an interruption, when is the miner trusted to be back?

5. Stable operating point

Accepted work returning is necessary evidence. It is not the last evidence.

Observe wall power, reported watts, observed hashrate, accepted work, rejects or stales when available, temperatures, fan behavior, board state, errors and resets across a declared window. Look for recurrence, oscillation and drift, not only a favorable end value.

The official VNISH fleet guidance gives the sequence clearly: baseline window, first boot check, early observation and an agreed evaluation window. It says the minimum interval is agreed per device cohort before the pilot. That is the right discipline here. Select a window that covers the site's relevant operating cycles, document interruptions, and explain why it is sufficient for this decision. Do not call any duration universal.

The hidden operational cost of a false recovery

A premature green light creates cost even when nobody writes a dollar figure beside it.

Equipment can be energized before useful accepted work fully returns. Technicians may revisit miners cleared too early. Repeated resets restart observation. A fleet may look recovered in aggregate while one cohort remains unstable. An unrecorded profile change can destroy the baseline comparison.

The cost also appears as uncertainty. Without first-acceptance time, the team cannot separate recovery from normal pool variance. Without wall measurement, it cannot establish the physical load. Without error timestamps, it cannot separate an old warning from recurrence.

None of this proves those costs are common or assigns them a universal price. It identifies the work that disappears when recovery is reduced to one status light.

The recovery window is the distance between powered and trusted.

The Restart Recovery Card

Use one card per canary or coherent restart cohort. Record fields only when they exist. Mark unavailable fields as unavailable.

FieldRecordWhy it matters
Event identityEvent ID, type, reason, start, ownerKeeps maintenance, network, power and curtailment events distinct
Device or cohortExact assets and cohort definitionPrevents unlike machines from sharing one recovery verdict
Hardware routeExact model and control-board routeAnchors the supported build path
Firmware stateCurrent build, profile, configured limit, change ownerShows what the device was asked to run
Pre-event baselineWall power, telemetry, pool fields, thermals, errors, windowGives recovery a real comparison point
Site contextPower-feed group, approved wall ceiling, cooling zone, network pathRecords physical and network constraints
Recovery authorizationWho approved re-energizing or reconnectingKeeps safety and operations ownership explicit
Device onlineTimestamp and exact system definitionRecords reachability without calling it recovery
First boot checkBuild, profile, boards, chips, warnings, timeCaptures early configuration and hardware state
Autotune stageInvoked or not, start, finish, result labelKeeps tuning separate from startup and stabilization
Wall powerInstrument, boundary, timestamp, value or rangeRecords physical demand at the named point
Firmware wattsExact interface label, timestamp, value or rangePreserves the device estimate without relabeling it
Pool connectionPool, endpoint, worker, first connection timeShows the work path is open
Submitted workExact field and first observed timeShows results are leaving the miner
Accepted workExact pool field, first acceptance, window total or seriesShows recognized work returned
Rejects and stalesExact pool fields when availableAdds delivery quality context
Operating signalsRequested and observed hashrate, temperatures, fans, boardsShows how the operating point develops
RecurrenceErrors, resets, missing hardware, timestampsTests whether the incident condition returns
Observation windowStart, end, timezone, cycles, interruptionsMakes the decision reproducible
Stop rulesSite-specific electrical, thermal, hardware, pool and stability rulesDefines when evidence overrides momentum
DecisionGO, HOLD or STOP, evidence and ownerCreates an auditable rollout gate

The card needs a consistent record across the tools that own each signal.

An eight-step staged restart

1. Classify the event

Name what changed. A network interruption may start with path and endpoint checks. A power event requires site-authorized electrical checks. Maintenance requires confirmation of the exact work performed. Curtailment return needs the operating plan and site event history preserved. For deeper curtailment context, read Bitcoin Mining Is Flexible, But Not at Any Price.

2. Freeze the baseline and stop rules

Preserve the last valid pre-event window when available. Declare site-specific stop rules, responsible people and measurement boundaries before restarting the pilot. Missing baseline evidence is not permission to invent it. Record the gap.

3. Form coherent restart cohorts

Group machines by the variables that can change this recovery decision, such as exact hardware route, build, power-feed group, cooling zone, maintenance state and observed condition. The VNISH Global cohort guide explains why a fleet average can hide the response tail.

4. Start one canary per cohort

Restart one representative canary before the wave. Keep an untouched peer or control group where the event and site design make that possible. Do not use a miner with unresolved faults as the only representative of a healthy cohort.

5. Confirm first boot before tuning

Verify identity, expected build, configured profile, visible boards and chips, warnings and recovery readiness. If autotune is part of the supported plan, record it as a separate stage. Do not mix its start or completion with the startup timestamp.

6. Rebuild the work path

Record pool connection, job flow when visible, submitted work, successful acceptance and error fields. Align the time zone and window across miner, pool and site logs. Do not treat first acceptance as proof of steady state.

7. Observe the operating point

Compare the canary with its baseline and untouched peer through the declared window. Keep the site-approved wall-power ceiling separate from the configured firmware setting. The VNISH Ninja watt-budget guide provides the measurement discipline for that boundary.

Watch the series, not one screenshot: wall power, reported watts, accepted work, hashrate, thermals, fan response, board state, errors and resets. Record site or network interruptions rather than smoothing them away.

8. Decide before the next wave

Choose GO, HOLD or STOP for that cohort. If GO, advance a controlled batch and repeat the gates. A canary clears a next wave, not the entire fleet forever. If the evidence splits, form a separate cohort rather than averaging away the difference.

GO, HOLD or STOP

DecisionUse whenNext actionIt does not mean
GOThe canary met prewritten site criteria across the declared recovery windowAdvance only its coherent cohort to the next controlled waveEvery machine, site or future restart is cleared
HOLDEvidence is incomplete, windows do not align, an interruption invalidated comparison, or behavior has not settledPreserve the current state, resolve the evidence gap and restart observation if appropriateThe firmware or hardware failed
STOPA prewritten electrical, thermal, hardware, pool-work, error, reset or stability condition was reachedContain the wave, preserve evidence and follow qualified support or recovery procedureOne universal threshold applies elsewhere or recovery is guaranteed

Persistent reset, board, thermal or accepted-work issues may move a machine out of the recovery workflow and into triage. The ROI ASIC tune, repair or replace decision guide helps structure that next question without assuming firmware caused the issue.

What this does not prove

A completed card does not prove that firmware caused every observed change. Power quality, cooling, machine condition, pool behavior, network path, maintenance and instrumentation can also move the result.

One recovery window does not forecast another site, season, event type, pool, build or hardware age. A canary reduces uncertainty for a defined cohort. It does not eliminate uncertainty.

An accepted share does not guarantee a future payment, stable hashrate, stable power or stable temperature. A stable window does not guarantee hardware life, revenue, efficiency, recovery success or freedom from another incident.

Aftermarket firmware and changed operating profiles alter machine operating conditions. They can affect power draw, temperature, stability, component stress, hardware life and manufacturer warranty or support conditions. Results depend on the exact model, control board, silicon, PSU, power input, cooling, network, pool, ambient conditions, build and settings. Qualified personnel must approve electrical work, limits and recovery procedures.

ROI ASIC is part of the VNISH ecosystem. Use the exact supported firmware route, preserve a recovery path appropriate to the board and site, and start with one measured canary. No performance, savings, uptime, recovery, revenue or financial result is guaranteed.

The next restart starts now

Do not wait for the next incident to invent the card.

Take one recent restart. Reconstruct the event type, online timestamp, first accepted work, wall measurement, error history and steady-state window. The missing fields will show what the next recovery plan needs.

Then choose one coherent cohort and one canary. Write the stop rules before the restart. Keep the evidence chain visible until GO, HOLD or STOP has an owner.

Your miner is not back because a tile changed color. It is back when the declared evidence says the operating point returned, useful work stayed accepted, and the next rollout decision is defensible.

Review six source-linked performance reports before converting a reported TH/s, watt or J/TH figure into an economic assumption. Open the independent field reports.

Sources and data freshness

  1. VNISH GLOBAL, VNISH 1.3.5 release page, current page checked 23 August 2026. Used for route identity and the distinction among installation, autotune and final stabilization. No release-date claim imported.
  2. VNISH GLOBAL, operational platform, fleet rollout guidance and data catalog, checked 23 August 2026. Demo values are examples, not performance evidence.
  3. Stratum V2, Mining Protocol specification, checked 23 August 2026. Used for the separation of connection, share submission, success and errors.
  4. NIST/SEMATECH, Quality versus reliability, checked 23 August 2026. Used only for the snapshot-versus-time distinction.
  5. BITMAIN, ANTMINER S21 specification, updated 8 April 2024 and checked 23 August 2026. Figures apply only to the exact base S21 and stated conditions.

Published by the ROI ASIC Analytics Desk. ROI ASIC is part of the VNISH ecosystem. This field guide does not promise a performance, uptime, recovery or financial result.