Keelix Systems
Menu
Contact

Why run a new system beside the current operation?

Parallel operation introduces a bounded new path beside a known fallback. It creates evidence about real operating behavior before the organization retires the path it already knows how to recover.

Answer

Run the new system beside a known fallback so the team can compare complete operating outcomes before transferring authority. Begin with explicit entry criteria, traffic allocation, stable identifiers, exception ownership, rollback, and evidence gates. Expand, restrict, or stop from those gates; do not keep both paths indefinitely.

A migration changes where the operation places authority. Replacing the current path as soon as the new system produces a credible result removes the comparison and recovery path at the moment they become most valuable. A bounded parallel period keeps the known operation available while the new one faces representative work.

Parallel operation is temporary evidence work, not duplicate work forever. Define what enters each path, what each path may change, how results are compared, who resolves exceptions, and what ends the period. Without those decisions, two systems can create two versions of operating truth.

Published Reviewed

Bound the operation and keep a known fallback.

Choose one unit of work and state where parallel operation begins and ends. The fallback should be a path the organization knows how to run and recover. It may be the current system, a manual procedure, or a narrower mode that preserves required records and decisions.

State what the new system may do in each mode. Shadow work changes no authoritative state. Supervised work may prepare an action for review. Authoritative work may complete specified actions. These labels matter only when the permitted actions and records are explicit.

Set entry criteria, allocation, and stable identity.

Define which cases qualify, which remain excluded, and how traffic moves between paths. Allocation may start with selected case types, a controlled sample, one operating team, or a scheduled window. Preserve exclusions for cases whose data, authority, or recovery path has not been evaluated.

Give each unit a stable identifier across both paths. Record the input version, system version, decisions, actions, timestamps, exceptions, and final state. Stable identity prevents duplicate work and lets the team reconcile differences without guessing whether two records describe the same case.

Compare end states, exceptions, and recovery.

A plausible answer is not an operating result. Compare whether the required record changed, the right owner received the exception, a message reached its destination, a transaction remained unique, and the item finished in a known state. Include time and review burden where they affect the operation.

Exercise dependency failure, retry, correction, rollback, and reconciliation. Name an owner for each exception class and give that person the source context needed to decide. Parallel operation earns confidence by exposing recovery behavior, not by producing two matching outputs on clean examples.

Use fixed evidence triggers to change authority.

Write the triggers before the run. Evidence may permit broader case coverage, a move from shadow to supervised action, or authority over one bounded step. The same record should identify conditions that restrict scope, return work to supervision, or stop the run.

Keep the decision with an accountable operational owner. The implementation team can assemble comparisons and explain technical causes, but it should not move the acceptance standard after seeing results. Record the decision, evidence reviewed, permitted boundary, and unresolved conditions.

End the parallel period with a recovery decision.

Before retiring the fallback, prove that operators can restrict authority, locate unfinished work, restore the fallback, and reconcile actions already taken. A paper rollback path is insufficient when the migration removes credentials, staff access, data flow, or procedural knowledge needed to use it.

Close parallel operation by choosing an authoritative path and documenting remaining supervised or excluded cases. If evidence does not support migration, narrow or stop the new path. Continuing two complete operations without a decision adds cost and increases the chance that their records diverge.

Choose the operating position.

Authority modes during parallel operation
Option Fits when Caution
Shadow The new path can process representative work without changing the authoritative operation. Compare complete end states and record what the shadow path could not observe.
Supervised The new path may propose or prepare work while a person controls material actions. Measure corrections, exceptions, review load, and recovery rather than output quality alone.
Authoritative Accepted evidence supports bounded action and the fallback has been exercised. Keep rollback and reconciliation available until the migration decision is complete.

Conditions that change readiness

  • The current operation provides a known fallback and recoverable state.
  • The new path has bounded entry criteria and traffic allocation.
  • Stable identifiers allow end-state comparison and reconciliation.
  • A named owner can expand, restrict, stop, or roll back authority.

Failure modes to test

  • Dual operation continues without an authority decision or end condition.
  • The team compares model outputs instead of complete operating end states.
  • The evidence gate moves after results are observed.
  • The fallback is retired before recovery and reconciliation are exercised.

Parallel-operation readiness

Set these conditions before representative work enters the new path or any authority moves from the fallback.

  • The unit of work and parallel-operation boundary are explicit.

  • Entry criteria, exclusions, and traffic allocation are recorded.

  • Shadow, supervised, and authoritative actions are defined.

  • Each unit keeps a stable identifier across comparison and retry.

  • The comparison method covers final records, exceptions, review, and recovery.

  • Exception classes have owners and usable decision context.

  • Expand, restrict, and stop triggers are written before results are reviewed.

  • A named owner makes and records each authority decision.

  • Rollback, unfinished-work discovery, and reconciliation have been exercised.

  • The parallel period has an end condition and a planned authoritative state.

Use parallel operation to make one accountable decision.

Running beside the current operation is useful when it reveals how the new system behaves under real conditions while preserving recovery. The comparison should make authority easier to place, not create a permanent second operation.

Finish with a recorded choice to expand, restrict, stop, or migrate. Keep the evidence, accepted boundary, and recovery record available for the next change in model, provider, workflow, or authority.

Reference record

Primary sources

  1. NIST AI RMF Playbook National Institute of Standards and Technology Accessed
  2. NIST AI 600-1: Generative Artificial Intelligence Profile National Institute of Standards and Technology Accessed

Apply the guide to a real operational system.

A short description of the workflow, system, or operating condition is enough to begin.