Files
AirCoding/AirPlan/docs/architecture/scheduler-state-machine-v1.md
AirCoding 82f3140847 Initial commit: AirCoding V1.0.0 Alpha architecture baseline
Complete architecture document set with multi-model review remediation:
- Frozen interface contracts, runtime semantics, DB schemas
- Event/tool/error/provider registries
- Scheduler and main agent state machines
- C4 module/code views, solution architecture, baseline V1
- Multi-model review reports and joint assessment
- Phase-gate remediation complete (P0/P1/P2/UX resolved)
- Implementation plan with T-000A through T-045
- Reference folders kept as placeholders only
2026-05-28 18:45:01 +08:00

486 lines
14 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# AirCoding Scheduler State Machine V1
Date: 2026-05-27
Status: Canonical Scheduler state machine for V1.0.0 Alpha skeleton
This document defines the MVP Scheduler lifecycle, task states, retry policy, workspace/merge flow, heartbeat handling, and escalation rules.
The Scheduler is not a coding agent. It is the orchestration service that reads TaskGraph state, launches worker agents, monitors progress, persists transitions, and escalates when architecture or user decisions are required.
## 1. Inputs and Outputs
Primary inputs:
```text
TaskSpec records
task_dependencies
RuntimeEvents
Architecture impact assessments
Permission decisions
WorkerResult records
Heartbeat/progress events
Resource availability snapshot
```
Primary outputs:
```text
task/agent/workspace events
child agent processes
workspace creation/merge requests
retry attempts
blocked/escalation reports
ProjectionStore updates via EventBus
```
The Scheduler persists all durable state through EventStore and domain tables. In-memory queues are rebuildable from SQLite.
Task dependencies use one edge model across TaskSpec, events, DB, and Scheduler memory:
```ts
interface TaskDependencySpec {
depends_on_task_id: TaskID
dependency_type: "hard" | "soft" | "conflict" | "serialization"
reason?: string
source?: "architecture" | "scheduler" | "worker" | "user" | "system"
}
```
Architecture Designer may create planned `hard`/`soft`/`serialization` edges. Scheduler may dynamically add `conflict` or `serialization` edges after write-area/resource analysis, but those inferred edges must be persisted through `task_dependencies` and durable events so restart recovery does not need to rediscover them.
## 2. Task Status Model
Task table status values:
```text
pending | running | completed | failed | blocked | cancelled | interrupted
```
State meaning:
| Status | Meaning | Scheduler action |
|---|---|---|
| `pending` | task exists but is not active | wait for dependencies/resources/write-set availability |
| `running` | task has an assigned agent/attempt | monitor heartbeat, events, timeout, result |
| `completed` | task met acceptance and produced WorkerResult | unblock dependents and consider merge/review |
| `failed` | task attempted and did not meet goal | retry, skip if allowed, or escalate |
| `blocked` | task cannot continue without higher-level decision | route to Main/Architecture Designer/user as needed |
| `cancelled` | task intentionally stopped | cancel dependents or replan |
| `interrupted` | task stopped but can be resumed or requeued | resume/requeue after requirement/control decision |
## 3. Task Attempt Model
Each run creates a `task_attempts` row.
```text
task.created
→ task.started + task_attempts.insert(attempt_index=N)
→ agent.started
→ terminal event:
task.completed
task.failed
task.blocked
task.cancelled
task.interrupted
```
Attempt indexes are monotonically increasing per task.
Failure signature is normalized from:
```text
error_kind + tool/command name + semantic diagnostic signature + failure location + concise failure summary
```
Repeated identical failure signatures consume retry budget faster than new failure signatures.
## 4. Scheduler Lifecycle
Top-level Scheduler states:
```text
IDLE
→ LOADING_GRAPH
→ PLANNING_WAVE
→ DISPATCHING
→ MONITORING
→ COLLECTING_RESULTS
→ MERGING
→ REVIEWING_WAVE
→ REPAIRING_OR_CONTINUING
→ COMPLETED | BLOCKED | CANCELLED
```
### `IDLE`
Scheduler has no active graph or all graphs are terminal.
Transitions:
- `task.created` / execution request → `LOADING_GRAPH`
- requirement change affecting active work → `LOADING_GRAPH`
### `LOADING_GRAPH`
Reads tasks, dependencies, attempts, agents, workspaces, and recent architecture/permission events from DB.
Actions:
1. Reconstruct TaskGraph.
2. Detect orphaned running tasks/agents.
3. Convert missing heartbeat/process to `agent.lost` and then `task.failed` or `task.interrupted`.
4. Validate dependency references.
5. Compute blocked/unblocked sets.
Transitions:
- valid runnable graph → `PLANNING_WAVE`
- no runnable tasks and blocked tasks exist → `BLOCKED`
- all terminal success → `COMPLETED`
- graph invalid beyond scheduler repair → `BLOCKED`
### `PLANNING_WAVE`
Computes next wave of tasks.
Inputs:
```text
hard dependencies
soft dependencies
conflict/serialization dependencies
write_area and expected_files
resource snapshot
model policy
permission profile
workspace strategy
```
Rules:
1. Hard dependencies must be completed before dispatch.
2. Soft dependencies influence priority/context but may run concurrently if Scheduler decides it is safe.
3. Tasks with different write areas may run concurrently.
4. Tasks with same file but provably different regions may run concurrently in separate worktrees.
5. Same write area with uncertain conflict serializes by default.
6. Reviewer tasks are read-only and may run concurrently with each other, but not against unstable unmerged write outputs unless they target a workspace snapshot.
7. Debugger tasks may serialize with tasks touching the same failure surface.
8. Machine resources constrain concurrency.
Wave output:
```ts
interface SchedulerWavePlan {
wave_id: string
runnable_task_ids: string[]
serialized_task_ids: string[]
workspace_assignments: Record<string, WorkspacePlan>
model_assignments: Record<string, ModelAssignment>
reason: string
}
```
Transitions:
- runnable wave exists → `DISPATCHING`
- no runnable wave but graph not terminal → `BLOCKED`
### `DISPATCHING`
Creates workspaces if needed, assembles context, and starts child agents.
Actions per task:
1. Create workspace (`main`, `worktree`, or `isolated_copy`).
2. Create `task.started` and `agent.started` events.
3. Write `agent.start` control message containing `TaskSpec`, `ContextPack`, and runtime context.
4. Record attempt row.
Dispatch failures:
- permission denied → `task.blocked`
- workspace create failure → retry if recoverable, otherwise `task.failed`
- context assembly failure → `task.failed` or `task.blocked` depending on reason
- process start failure → `task.failed`
Transition: `MONITORING`.
### `MONITORING`
Tracks active agents.
Inputs:
```text
agent.heartbeat
task.progress
tool/command events
process exit
soft timeout
hard timeout
user interruption
requirement.changed
```
Heartbeat policy:
```text
agent emits heartbeat periodically
Scheduler updates agents.last_heartbeat_at and tasks.heartbeat_at
missing heartbeat past threshold → inspect process
process alive but silent → send status ping or soft cancel depending on timeout
process gone without result → agent.lost
```
Timeout policy:
| Timeout | Action |
|---|---|
| soft timeout | ask agent for checkpoint/status; Scheduler may extend if progress is credible |
| hard timeout | cancel/kill worker; emit `agent.lost` or `agent.cancelled`; mark task failed/interrupted |
Requirement change policy:
1. Main Agent records `requirement.changed`.
2. Scheduler pauses affected dispatch decisions.
3. Architecture Designer assesses impact when change may affect interface/architecture/product goals.
4. Implementation-only non-conflicting change may be silently absorbed.
5. Architecture-level impact uses `architecture.impact.completed` and may require user confirmation or replan.
Transitions:
- all active agents terminal → `COLLECTING_RESULTS`
- user/global cancellation → `CANCELLED`
- architecture/user blocker → `BLOCKED`
### `COLLECTING_RESULTS`
Consumes WorkerResult and terminal task events.
Actions:
1. Validate WorkerResult against output contract.
2. Persist artifacts/evidence references.
3. Mark attempt terminal.
4. Classify task terminal status.
5. Compute dependency unblocks.
Status mapping:
| WorkerResult status | Task status |
|---|---|
| `completed` | `completed` |
| `failed` | `failed` pending retry/skip decision |
| `blocked` | `blocked` |
| `cancelled` | `cancelled` or `interrupted` depending on resumability |
Transition:
- write workspaces need merge → `MERGING`
- no merge needed but review required → `REVIEWING_WAVE`
- more work remains → `REPAIRING_OR_CONTINUING`
- all done → `COMPLETED`
### `MERGING`
Merges completed workspace outputs back into the target workspace.
Strategies:
```text
main → no merge
worktree → git merge or patch apply
isolated_copy → copy-back or patch apply
```
Rules:
1. Merge only tasks with successful WorkerResult unless explicitly recovering artifacts.
2. Serialize merges touching overlapping write areas.
3. Persist `workspace.merge.started` and terminal merge event.
4. Conflict creates `workspace.merge.conflicted` and either schedules Debugger/repair or blocks.
Conflict handling:
| Conflict | Scheduler action |
|---|---|
| trivial patch conflict | schedule repair/debugger or retry serially |
| semantic conflict | route to Reviewer/Architecture Designer |
| architecture/interface conflict | block and request Architecture impact assessment |
| user requirement conflict | block and ask Main Agent/user |
Transition:
- merge success and review needed → `REVIEWING_WAVE`
- merge success and more tasks → `REPAIRING_OR_CONTINUING`
- merge conflict recoverable → `REPAIRING_OR_CONTINUING`
- unrecoverable conflict → `BLOCKED`
### `REVIEWING_WAVE`
Schedules review tasks after implementation/debug waves when required by plan or risk.
Reviewers are read-only unless a follow-up execute/debug task is created.
Review result handling:
```text
approved → continue
minor issue with clear fix → create execute repair task
high-severity/security/architecture issue → block or Architecture Designer assessment
insufficient evidence → create verification task
```
Transition: `REPAIRING_OR_CONTINUING`.
### `REPAIRING_OR_CONTINUING`
Decides next action after a wave.
Actions:
1. Retry failed tasks within budget.
2. Create repair tasks from review/debug output.
3. Skip allowed failed tasks only when acceptance criteria permit fallback.
4. Recompute graph and continue.
5. Escalate if architecture/user decision is required.
Transition:
- more runnable tasks → `PLANNING_WAVE`
- blocked → `BLOCKED`
- all acceptance criteria met → `COMPLETED`
- cancelled → `CANCELLED`
## 5. Retry Policy
Default retry budget: task-specific `TaskSpec.constraints.retry_budget`, usually 35.
Retry rules:
1. Retry only when failure is plausibly recoverable.
2. First retry may use same agent strategy with corrected context.
3. Later retries should change one dimension: model, context, command strategy, serialization, or Debugger involvement.
4. Identical failure signature after multiple attempts triggers escalation faster.
5. Environment/kernel/toolchain impossibility becomes `blocked`, not infinite retry.
6. Architecture/interface mismatch routes to Architecture Designer.
7. If fallback is allowed and the failed task is non-critical, Scheduler may skip after recording evidence and risk.
Retry decision shape:
```ts
interface RetryDecision {
action: "retry" | "retry_serial" | "debug" | "skip" | "block" | "cancel"
reason: string
next_model_policy?: "scheduler_forced" | "agent_select"
required_context_refs?: string[]
}
```
## 6. Model Assignment Policy
Scheduler chooses per task:
```text
scheduler_forced → exact model assigned by Scheduler
agent_select → agent may choose within allowed model/capability constraints
```
Guidelines:
- Implementation tasks with hidden complexity should use strong coding-capable model assignment or constrained agent choice.
- Cross-model review should be scheduler-forced.
- Debug tasks may switch models after repeated failures.
- Cheap/simple tasks may use lower-cost models when quality risk is low.
## 7. Resource-Aware Concurrency
Scheduler considers:
```text
CPU cores
RAM pressure
disk space
active command count
LLM provider rate limits
toolchain locks
workspace write conflicts
user-configured max concurrency
```
Concurrency can shrink or expand dynamically. Running tasks are not killed solely because resources become tighter unless the system is at risk; new dispatch is paused first.
## 8. Escalation Model Hooks
Scheduler does not ask the user directly except through Main Agent/PermissionEngine.
Escalation routes:
| Trigger | Route |
|---|---|
| product/user choice needed | Main Agent |
| architecture/interface change | Architecture Designer → Main Agent/user if needed |
| permission decision | PermissionEngine → Main Agent/UI |
| environment impossible | Main Agent with Doctor/Debugger evidence |
| repeated implementation failure | Debugger or Main Agent based on recoverability |
| high-severity review finding | Main Agent + Architecture Designer if design-level |
## 9. Recovery on Restart
On startup/resume:
1. Load `tasks` where status is `running` or `interrupted`.
2. Load active `agents` and `workspaces`.
3. Check process liveness where PID is known.
4. If process is alive and IPC can reconnect, resume monitoring.
5. If process is gone, emit `agent.lost`; mark task failed/interrupted based on resumability.
6. Preserve workspaces until merge/cleanup decision is recorded.
7. Rebuild queue from pending/failed-with-retry tasks.
## 10. Terminal Graph States
### `COMPLETED`
All required tasks completed and required review/verification gates passed or documented as skipped by policy.
### `BLOCKED`
No runnable work remains because a decision or external condition is required. Scheduler must produce a concise blocker report with:
```text
blocked task IDs
reason
required decision
available evidence
safe suggested next step
```
### `CANCELLED`
User/system cancellation propagated to running agents, task statuses terminalized, and workspaces left in a recoverable state.
## 11. V1.0.0 Alpha Cut Line
The V1.0.0 Alpha skeleton must implement:
1. DB-backed TaskGraph loading.
2. Hard/soft dependency handling.
3. Write-area serialization and basic worktree workspace strategy.
4. Child process dispatch via NDJSON IPC.
5. Heartbeat timeout detection.
6. Retry budget and failure signature tracking.
7. Workspace merge success/conflict events.
8. Review wave support.
9. Restart recovery for running/lost agents.
10. Escalation report generation.
Advanced optimization can wait:
```text
fine-grained AST write-set prediction
multi-machine scheduling
provider rate-limit optimization
automatic semantic merge resolution
learning-based retry strategy
```