Skip to main content

RACS — Risk-Aware Coordination System

Risk-aware coordination for autonomous systems

RACS is an open coordination layer for detecting emerging operational risk and coordinating autonomous systems before local disruptions propagate across the broader system.

  • Open source
  • Observable telemetry only
  • 360 simulations
  • 248 automated tests
  • Reproducible V1

Why RACS

Autonomy creates a coordination problem

  • Autonomous systems increasingly operate as fleets and shared workflows.
  • Local degradation can affect queues, shared resources, and downstream processes.
  • Conventional recovery often begins after hard failure.
  • RACS explores whether emerging operational risk can be detected and coordinated earlier.
  1. 01Localized degradation
  2. 02Delayed work
  3. 03Queue pressure
  4. 04Downstream disruption

Reactive recovery

  1. 01Hard failure
  2. 02Detect
  3. 03Recover

RACS

  1. 01Observable degradation evidence
  2. 02Risk signal
  3. 03Coordinate
  4. 04Contain/recover

How RACS works

A separate coordination layer for operational risk

RACS sits beside autonomous operations as a coordination layer, converting observable system behavior into risk signals that can inform bounded coordination actions.

RACS system architecture

Sense

  1. Observable telemetry
  2. Risk prediction
  3. RiskSignal

Coordinate

  1. NetworkBrain

Guard

  1. SafetyGate

Act

  1. Coordination action
  2. Autonomous system

Observable telemetry to Risk prediction to RiskSignal to NetworkBrain to SafetyGate to Coordination action to Autonomous system

Coordination actions in V1

  • Predictive drain
  • Cordon
  • Hard-failure fallback

V1 experiment

A controlled test of predictive coordination

  • Multiple AMRs transport work.
  • Completed transport feeds a downstream workstation.
  • One AMR progressively loses service capacity.
  • Workload timing varies by seeded arrival sequence.

V1 operational flow

  1. 01AMRs
  2. 02Transport tasks
  3. 03Workstation buffer
  4. 04Downstream processing

Reference counterfactual

Healthy

No degradation is applied, providing a paired reference for the same workload realization.

Hard-failure recovery

Reactive baseline

Robot remains available until hard failure; stranded work is then recovered and reassigned.

Predictive coordination

RACS

Observable task behavior can trigger a predictive drain before hard failure. Existing work may complete while new assignments are redirected.

Study scale

30

paired workload seeds

4

disturbance-time WIP conditions

3

conditions per seed/WIP pair

360

total simulations

248

automated tests

Key results

What V1 found

Lower queue AUC and lower latency are favorable.

H1 — Transport coordination

RACS improved mean queue AUC and mean task latency across the seeded study, while individual workload realizations included regressions.

PARTIALLY SUPPORTED
Mean queue AUC delta, RACS − baseline
-2.9333
Mean latency delta, RACS − baseline
-0.0901
Queue AUC regressions
24 / 120 paired comparisons

H2 — Earlier intervention and downstream benefit

Downstream benefit was not greater in runs where RACS intervened before propagation than in runs where intervention occurred at or after propagation.

NOT ESTABLISHED

H3 — Intervention headroom

Within the tested WIP levels, greater disturbance-time WIP was associated with a higher fraction of runs in which RACS intervened before propagation.

SUPPORTED WITHIN TESTED WIP LEVELS
WIP 0
66.7%
WIP 1
80.0%
WIP 2
93.3%
WIP 4
93.3%

V1 EVIDENCE

Results from the seeded robustness study

View full V1 evaluation

Intervention before propagation by WIP

Higher disturbance-time WIP was associated with more frequent intervention before baseline propagation within the tested WIP levels.

Bar chart showing intervention-before-propagation fractions of 66.7 percent at WIP 0, 80.0 percent at WIP 1, and 93.3 percent at WIP 2 and WIP 4.
Fraction of paired RACS runs where intervention occurred before the paired reactive baseline's degradation-induced cascade start.

Queue AUC delta distribution

Mean queue AUC favored RACS, while 24 of 120 paired comparisons worsened.

Distribution plot of queue AUC deltas for RACS minus baseline, with some positive regression cases.
Paired queue AUC deltas use RACS minus reactive baseline; negative values favor RACS.

Baseline cascade timing by WIP

Baseline propagation timing varied by seeded arrivals and shifted later at higher disturbance-time WIP.

Distribution plot showing baseline cascade start timing by WIP level, with later timing at higher WIP.
Reactive-baseline degradation-induced cascade start distributions for WIP 0, 1, 2, and 4.

Representative timeline

The representative WIP 2 case shows the event sequence available in the frozen result artifact.

Timeline plot for the representative seed 1012 at WIP 2 showing degradation, intervention, and propagation events.
Representative case selected as the seed closest to median baseline queue AUC at WIP 2, breaking ties by lowest seed.

Limitations

What V1 does not establish

These limitations define the current scientific scope of the V1 evidence and separate completed work from future validation.

  • Discrete-event simulation
  • No physical-robot validation
  • Single-site scenario
  • One degrading AMR
  • Abstract downstream workstation
  • Deterministic robot service duration
  • Only workload-arrival timing randomized
  • 30 seeds with descriptive statistics
  • No statistical-significance claim
  • No ROS 2 / Gazebo validation
  • Downstream effect sizes remain small

Roadmap

Building toward broader autonomous-system resilience

The roadmap distinguishes completed V1 evidence from next-step evaluation and exploratory directions.

  1. V1COMPLETE

    Reproducible simulation and seeded robustness evaluation

  2. V1.5NEXT

    External evaluation and independent reproduction

  3. V2PLANNED

    Robotics middleware / simulator integration

  4. FutureEXPLORATORY

    Heterogeneous and broader operational coordination

Evaluate RACS

Evaluate RACS

We are inviting robotics engineers, researchers, and autonomous-systems practitioners to reproduce, critique, and extend the V1 experiment.