Spaces:
Sleeping
Sleeping
Strengthen grading transparency and reproducibility evidence
Browse files- README.md +214 -99
- WEB_README.md +56 -129
- scripts/reproducibility_check.py +125 -0
- server/cloud_devops_env_environment.py +79 -17
README.md
CHANGED
|
@@ -11,158 +11,273 @@ tags:
|
|
| 11 |
|
| 12 |
# Cloud DevOps RLEnv
|
| 13 |
|
| 14 |
-
Cloud DevOps RLEnv is an OpenEnv
|
| 15 |
|
| 16 |
-
|
| 17 |
|
| 18 |
## Judge-Aligned Snapshot
|
| 19 |
|
| 20 |
-
This section maps the environment directly to the scoring rubric.
|
| 21 |
-
|
| 22 |
| Parameter | Weight | How this environment addresses it |
|
| 23 |
| --- | --- | --- |
|
| 24 |
-
| Real-world utility | 30% |
|
| 25 |
-
| Task
|
| 26 |
-
| Environment design | 20% | Typed
|
| 27 |
-
| Code quality
|
| 28 |
-
| Creativity
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 29 |
|
| 30 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 31 |
|
| 32 |
-
|
| 33 |
|
| 34 |
-
-
|
| 35 |
-
-
|
| 36 |
-
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
|
| 38 |
-
|
| 39 |
|
| 40 |
-
|
|
|
|
|
|
|
| 41 |
|
| 42 |
-
|
| 43 |
-
- Ambiguous telemetry: logs expose symptoms and IPs, not always direct resource names.
|
| 44 |
-
- Action-cost pressure: every action incurs a negative reward, penalizing brute-force behavior.
|
| 45 |
-
- Multi-hop dependency: agent must use metadata resolution before remediation in medium/hard.
|
| 46 |
-
- Cascading failures: unresolved hard incidents degrade additional components after step 8.
|
| 47 |
|
| 48 |
-
|
| 49 |
|
| 50 |
-
|
|
| 51 |
| --- | --- | --- | --- |
|
| 52 |
-
|
|
| 53 |
-
|
|
| 54 |
-
|
|
| 55 |
|
| 56 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 57 |
|
| 58 |
-
|
| 59 |
-
- Deterministic transitions: action handling is rule-based and reproducible.
|
| 60 |
-
- Deterministic rewards: shaped by explicit achievement checkpoints and safety penalties.
|
| 61 |
-
- Deterministic success: inferred from resolved incident state, not ad hoc heuristics.
|
| 62 |
|
| 63 |
-
|
| 64 |
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68 |
|
| 69 |
-
|
| 70 |
|
| 71 |
-
|
| 72 |
-
-
|
| 73 |
-
|
| 74 |
-
|
|
|
|
|
|
|
| 75 |
|
| 76 |
-
##
|
| 77 |
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
|
|
|
| 83 |
|
| 84 |
-
|
|
|
|
| 85 |
|
| 86 |
-
|
| 87 |
-
- action must be allow or deny.
|
| 88 |
-
- Invalid actions are rejected and do not mutate state.
|
| 89 |
|
| 90 |
-
###
|
| 91 |
|
| 92 |
-
|
|
|
|
| 93 |
|
| 94 |
-
|
|
|
|
| 95 |
|
| 96 |
-
|
|
|
|
|
|
|
|
|
|
| 97 |
|
| 98 |
-
|
|
|
|
|
|
|
| 99 |
|
| 100 |
-
|
| 101 |
-
- MODEL_NAME
|
| 102 |
-
- HF_TOKEN
|
| 103 |
|
| 104 |
-
|
|
|
|
| 105 |
|
| 106 |
-
|
| 107 |
-
-
|
| 108 |
-
- Emits strict stdout contract only:
|
| 109 |
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 114 |
```
|
| 115 |
|
| 116 |
-
##
|
| 117 |
|
| 118 |
-
|
|
|
|
| 119 |
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
| easy | 2 | 0.780 | success=true |
|
| 123 |
-
| medium | 4 | 0.960 | success=true |
|
| 124 |
-
| hard | 5 | 0.999 | success=true |
|
| 125 |
|
| 126 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 127 |
|
| 128 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 129 |
|
| 130 |
-
|
| 131 |
-
-
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
|
| 135 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 136 |
|
| 137 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 138 |
|
| 139 |
From repository root:
|
| 140 |
|
| 141 |
```bash
|
| 142 |
-
#
|
| 143 |
..\\.venv\\Scripts\\openenv validate
|
| 144 |
|
| 145 |
-
#
|
| 146 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 147 |
|
| 148 |
-
|
| 149 |
-
bash scripts/validate-submission.sh https://<your-space>.hf.space .
|
| 150 |
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
docker run --rm -p 8000:8000 cloud-devops-env:latest
|
| 154 |
```
|
| 155 |
|
| 156 |
-
## Hugging Face Space Deployment
|
| 157 |
|
| 158 |
-
1. Keep front matter intact
|
| 159 |
-
2. Push
|
| 160 |
-
3. Configure
|
| 161 |
-
|
| 162 |
-
-
|
| 163 |
-
-
|
| 164 |
-
|
|
|
|
|
|
|
|
|
|
| 165 |
|
| 166 |
Reference:
|
| 167 |
-
|
| 168 |
- https://huggingface.co/docs/hub/spaces-config-reference
|
|
|
|
| 11 |
|
| 12 |
# Cloud DevOps RLEnv
|
| 13 |
|
| 14 |
+
Cloud DevOps RLEnv is an OpenEnv-compatible cloud incident-response benchmark designed for agentic SRE and DevOps workflows.
|
| 15 |
|
| 16 |
+
This environment rewards correct diagnosis and safe remediation, not blind action execution. It is deterministic, reproducible, and optimized for hackathon evaluation.
|
| 17 |
|
| 18 |
## Judge-Aligned Snapshot
|
| 19 |
|
|
|
|
|
|
|
| 20 |
| Parameter | Weight | How this environment addresses it |
|
| 21 |
| --- | --- | --- |
|
| 22 |
+
| Real-world utility | 30% | Models practical SRE outage response loops: telemetry triage, dependency mapping, and safe remediation. |
|
| 23 |
+
| Task & grader quality | 25% | Three deterministic tasks with explicit objectives, strict success gates, and reproducible scoring behavior. |
|
| 24 |
+
| Environment design | 20% | Typed action/observation/state models, clean reset semantics, shaped rewards, action-cost efficiency pressure, clear boundaries. |
|
| 25 |
+
| Code quality & spec compliance | 15% | OpenEnv-compliant project layout, Dockerized runtime, strict inference output contract, validation scripts. |
|
| 26 |
+
| Creativity & novelty | 10% | Multi-hop metadata dependency, cascading failures, high-decoy search space, and safety-aware penalties. |
|
| 27 |
+
|
| 28 |
+
## Why This Environment
|
| 29 |
+
|
| 30 |
+
Real incidents are multi-step and noisy. Good agents must:
|
| 31 |
+
- gather context before changing systems
|
| 32 |
+
- identify root cause from logs and topology
|
| 33 |
+
- apply minimal, correct fixes
|
| 34 |
+
- verify resolution
|
| 35 |
+
|
| 36 |
+
Cloud DevOps RLEnv simulates that behavior with realistic failure patterns, decoy resources, shaped rewards, and anti-shortcut guardrails.
|
| 37 |
+
|
| 38 |
+
## Why It's Hard
|
| 39 |
+
|
| 40 |
+
This benchmark is intentionally designed to resist brute-force policies and reward disciplined SRE reasoning:
|
| 41 |
|
| 42 |
+
- Needle-in-a-haystack discovery: 20+ decoy compute nodes and 20+ decoy security groups increase search complexity.
|
| 43 |
+
- Ambiguous telemetry: noisy, raw operational logs surface symptoms (including IP-only clues) rather than direct root-cause labels.
|
| 44 |
+
- Action-penalty heuristics: every action has a small negative cost, so efficient remediation beats command spamming.
|
| 45 |
+
- Multi-hop dependency resolution: agents must map IP addresses to resource IDs via metadata lookup before applying fixes.
|
| 46 |
+
- System drift under pressure: in hard mode, delayed remediation triggers cascading failures that worsen observability and reward dynamics.
|
| 47 |
|
| 48 |
+
## Environment Scope
|
| 49 |
|
| 50 |
+
- Domain: Cloud SRE / DevOps incident response
|
| 51 |
+
- Difficulty tiers: easy, medium, hard
|
| 52 |
+
- Max environment steps per episode: 20
|
| 53 |
+
- Runtime health states: CRITICAL, DEGRADED, HEALTHY
|
| 54 |
+
- Decoy resources: 20 backend instances + 20 backend security groups
|
| 55 |
+
|
| 56 |
+
## OpenEnv Compliance
|
| 57 |
+
|
| 58 |
+
Core files:
|
| 59 |
+
- openenv.yaml
|
| 60 |
+
- env.py
|
| 61 |
+
- models.py
|
| 62 |
+
- inference.py
|
| 63 |
+
- server/app.py
|
| 64 |
+
- server/cloud_devops_env_environment.py
|
| 65 |
|
| 66 |
+
Validator command:
|
| 67 |
|
| 68 |
+
```bash
|
| 69 |
+
..\\.venv\\Scripts\\openenv validate
|
| 70 |
+
```
|
| 71 |
|
| 72 |
+
## Action Space
|
|
|
|
|
|
|
|
|
|
|
|
|
| 73 |
|
| 74 |
+
Model: CloudAction
|
| 75 |
|
| 76 |
+
| Field | Type | Required | Description |
|
| 77 |
| --- | --- | --- | --- |
|
| 78 |
+
| command | enum | yes | One of: list_resources, describe_resource, view_logs, query_metadata, update_security_group, restart_service, submit_solution |
|
| 79 |
+
| resource_id | string | conditional | Required for most actions except list_resources and query_metadata |
|
| 80 |
+
| parameters | object | conditional | Used by mutating actions (for example, security-group updates) |
|
| 81 |
|
| 82 |
+
Action semantics:
|
| 83 |
+
- list_resources: Enumerates available resources including decoys.
|
| 84 |
+
- describe_resource: Returns structured details for one resource.
|
| 85 |
+
- view_logs: Returns logs for one resource.
|
| 86 |
+
- query_metadata: Resolves infrastructure metadata (for example, IP address to resource ID).
|
| 87 |
+
- update_security_group: Appends a rule (requires parameters.port and parameters.action where action is allow/deny).
|
| 88 |
+
- restart_service: Restarts one instance/service by ID.
|
| 89 |
+
- submit_solution: Declares the episode solved (or not solved).
|
| 90 |
|
| 91 |
+
## Observation And State Space
|
|
|
|
|
|
|
|
|
|
| 92 |
|
| 93 |
+
Observation model: CloudObservation
|
| 94 |
|
| 95 |
+
| Field | Description |
|
| 96 |
+
| --- | --- |
|
| 97 |
+
| output | Main command output |
|
| 98 |
+
| error | Error string for failed commands |
|
| 99 |
+
| system_health_status | CRITICAL, DEGRADED, HEALTHY |
|
| 100 |
+
| done | Episode terminal flag |
|
| 101 |
+
| reward | Step reward |
|
| 102 |
+
| metadata | Diagnostics such as task, step_count, resolved, achievements |
|
| 103 |
|
| 104 |
+
Hidden state model: CloudState
|
| 105 |
|
| 106 |
+
| Field | Description |
|
| 107 |
+
| --- | --- |
|
| 108 |
+
| task_difficulty | easy, medium, hard |
|
| 109 |
+
| resources | Full resource graph including logs/rules |
|
| 110 |
+
| step_count | Current step counter |
|
| 111 |
+
| is_resolved | Whether root cause has been fixed |
|
| 112 |
|
| 113 |
+
## Reward Design
|
| 114 |
|
| 115 |
+
Reward shaping is sparse-but-guided:
|
| 116 |
+
- discovery rewards for correct investigative steps
|
| 117 |
+
- larger terminal rewards for correct remediation
|
| 118 |
+
- penalties for unsafe or premature operations
|
| 119 |
+
- fixed action cost per step (efficiency pressure)
|
| 120 |
+
- timeout terminal condition after max steps
|
| 121 |
|
| 122 |
+
Per-step reward is clipped to [-1.0, 1.0].
|
| 123 |
+
Inference task score is adjusted to remain strictly within (0.0, 1.0) for Phase-2 validator compatibility.
|
| 124 |
|
| 125 |
+
## Detailed Task Playbooks
|
|
|
|
|
|
|
| 126 |
|
| 127 |
+
### Easy Task
|
| 128 |
|
| 129 |
+
Incident:
|
| 130 |
+
- Web traffic blocked by security group.
|
| 131 |
|
| 132 |
+
Objective:
|
| 133 |
+
- Open port 80 on sg-web.
|
| 134 |
|
| 135 |
+
Typical successful sequence:
|
| 136 |
+
1. list_resources
|
| 137 |
+
2. describe_resource(sg-web) for context (+0.2)
|
| 138 |
+
3. update_security_group(sg-web, port=80, action=allow) (+0.8, done)
|
| 139 |
|
| 140 |
+
Expected score:
|
| 141 |
+
- ~0.97 for full playbook with efficient triage
|
| 142 |
+
- ~0.79 if agent skips the optional read step
|
| 143 |
|
| 144 |
+
### Medium Task
|
|
|
|
|
|
|
| 145 |
|
| 146 |
+
Incident:
|
| 147 |
+
- API cannot reach DB due to blocked port 5432.
|
| 148 |
|
| 149 |
+
Objective:
|
| 150 |
+
- Confirm root cause from logs, then open port 5432 on sg-db.
|
|
|
|
| 151 |
|
| 152 |
+
Typical successful sequence:
|
| 153 |
+
1. list_resources
|
| 154 |
+
2. view_logs(i-api) to identify DB timeout (+0.2)
|
| 155 |
+
3. query_metadata(ip_address=10.0.4.5) to resolve DB target (+0.2)
|
| 156 |
+
4. update_security_group(sg-db, port=5432, action=allow) (+0.6, done if logs and metadata lookup were completed)
|
| 157 |
+
|
| 158 |
+
Guardrail:
|
| 159 |
+
- Applying the SG change before log triage + metadata lookup gives a penalty (-0.1) and does not close the incident.
|
| 160 |
+
|
| 161 |
+
Expected score:
|
| 162 |
+
- ~0.97 with full investigative path (logs -> metadata lookup -> remediation)
|
| 163 |
+
- below ~0.90 when metadata dependency is skipped
|
| 164 |
+
|
| 165 |
+
## Determinism And Grader Transparency
|
| 166 |
+
|
| 167 |
+
- Deterministic reset and transitions: no randomization is used in task generation or transition logic.
|
| 168 |
+
- Transparent grading signals: observation metadata includes achievements, resolution status, termination reason, and reward breakdown events.
|
| 169 |
+
- Reproducibility helper:
|
| 170 |
+
|
| 171 |
+
```bash
|
| 172 |
+
..\\.venv\\Scripts\\python scripts/reproducibility_check.py
|
| 173 |
```
|
| 174 |
|
| 175 |
+
### Hard Task
|
| 176 |
|
| 177 |
+
Incident:
|
| 178 |
+
- Checkout path degraded due to upstream timeout to an IP-only target that must be resolved first.
|
| 179 |
|
| 180 |
+
Objective:
|
| 181 |
+
- Trace LB errors to the correct target, resolve resource identity via metadata, and restart i-web2 only after diagnosis.
|
|
|
|
|
|
|
|
|
|
| 182 |
|
| 183 |
+
Typical successful sequence:
|
| 184 |
+
1. list_resources
|
| 185 |
+
2. view_logs(lb-main) to identify failing upstream IP (+0.2)
|
| 186 |
+
3. query_metadata(ip_address=<failing_ip>) to resolve target ID (+0.2)
|
| 187 |
+
4. describe_resource(i-web2) or view_logs(i-web2) (+0.2)
|
| 188 |
+
5. restart_service(i-web2) (+0.8, done when all investigation achievements exist)
|
| 189 |
|
| 190 |
+
Guardrails:
|
| 191 |
+
- Restarting i-web2 before investigation: penalty (-0.1), no resolution.
|
| 192 |
+
- Restarting healthy i-web1: penalty (-0.2).
|
| 193 |
+
- Premature submit_solution in hard mode: penalty (-0.1), episode continues.
|
| 194 |
+
- If unresolved after step 8 in hard mode, lb-external also fails (cascading failure), increasing pressure and noise.
|
| 195 |
|
| 196 |
+
Expected score:
|
| 197 |
+
- near 1.0 after score clamping for strong trajectories (can exceed 1.0 raw before clamp)
|
| 198 |
+
|
| 199 |
+
## API Endpoints
|
| 200 |
+
|
| 201 |
+
Core runtime:
|
| 202 |
+
- GET /health
|
| 203 |
+
- POST /reset
|
| 204 |
+
- POST /step
|
| 205 |
+
- GET /state
|
| 206 |
+
- GET /schema
|
| 207 |
+
- WS /ws
|
| 208 |
+
|
| 209 |
+
Web UI runtime:
|
| 210 |
+
- GET /web
|
| 211 |
+
- POST /web/reset
|
| 212 |
+
- POST /web/step
|
| 213 |
+
- GET /web/state
|
| 214 |
+
- GET /web/metadata
|
| 215 |
+
|
| 216 |
+
## Inference Contract
|
| 217 |
+
|
| 218 |
+
inference.py requirements:
|
| 219 |
+
- uses OpenAI client
|
| 220 |
+
- reads API_BASE_URL, MODEL_NAME, HF_TOKEN
|
| 221 |
+
- emits strict logs: [START], [STEP], [END]
|
| 222 |
|
| 223 |
+
Current defaults in code:
|
| 224 |
+
- MODEL_NAME default: google/gemma-4-26B-A4B-it
|
| 225 |
+
- MAX_STEPS (in inference loop): 15
|
| 226 |
+
- success flag is derived from environment resolution state
|
| 227 |
+
|
| 228 |
+
## Baselines
|
| 229 |
+
|
| 230 |
+
### Deterministic task baselines
|
| 231 |
+
|
| 232 |
+
| Task | Typical baseline score |
|
| 233 |
+
| --- | --- |
|
| 234 |
+
| easy | 0.78 |
|
| 235 |
+
| medium | 0.96 |
|
| 236 |
+
| hard | 0.999 |
|
| 237 |
+
|
| 238 |
+
### LLM policy comparison
|
| 239 |
+
|
| 240 |
+
| Model | Easy | Medium | Summary |
|
| 241 |
+
| --- | --- | --- | --- |
|
| 242 |
+
| gemma-3-27b-it | 0.2 | 0.2 | Underperformed on this environment |
|
| 243 |
+
| gemma-4-31b-it | 1.0 | 1.0 | Perfect on both easy and medium |
|
| 244 |
+
|
| 245 |
+
## Local Setup And Validation
|
| 246 |
|
| 247 |
From repository root:
|
| 248 |
|
| 249 |
```bash
|
| 250 |
+
# Structure + manifest validation
|
| 251 |
..\\.venv\\Scripts\\openenv validate
|
| 252 |
|
| 253 |
+
# Determinism/reproducibility smoke test
|
| 254 |
+
..\\.venv\\Scripts\\python scripts/reproducibility_check.py
|
| 255 |
+
|
| 256 |
+
# Submission-oriented local checks (without live inference)
|
| 257 |
+
bash scripts/pre_submit_validate.sh --skip-inference
|
| 258 |
+
|
| 259 |
+
# Build local image
|
| 260 |
+
docker build -t cloud-devops-env:phase1 -f Dockerfile .
|
| 261 |
+
```
|
| 262 |
|
| 263 |
+
Optional local server:
|
|
|
|
| 264 |
|
| 265 |
+
```bash
|
| 266 |
+
uvicorn server.app:app --host 0.0.0.0 --port 8000
|
|
|
|
| 267 |
```
|
| 268 |
|
| 269 |
+
## Hugging Face Space Deployment
|
| 270 |
|
| 271 |
+
1. Keep this front matter block intact (includes mandatory openenv tag).
|
| 272 |
+
2. Push to Space (Docker SDK).
|
| 273 |
+
3. Configure secrets/variables:
|
| 274 |
+
- HF_TOKEN
|
| 275 |
+
- API_BASE_URL (for example https://router.huggingface.co/v1)
|
| 276 |
+
- MODEL_NAME
|
| 277 |
+
4. Wait for build completion.
|
| 278 |
+
5. Verify:
|
| 279 |
+
- GET /health returns 200
|
| 280 |
+
- POST /reset returns 200
|
| 281 |
|
| 282 |
Reference:
|
|
|
|
| 283 |
- https://huggingface.co/docs/hub/spaces-config-reference
|
WEB_README.md
CHANGED
|
@@ -1,172 +1,99 @@
|
|
| 1 |
-
# Cloud DevOps RLEnv
|
| 2 |
|
| 3 |
-
Use this page as
|
| 4 |
|
| 5 |
-
##
|
| 6 |
-
|
| 7 |
-
You are the on-call SRE for a production outage.
|
| 8 |
-
|
| 9 |
-
Your loop is always:
|
| 10 |
|
| 11 |
1. Triage quickly.
|
| 12 |
-
2.
|
| 13 |
-
3. Apply the minimum safe
|
| 14 |
-
4. Verify
|
| 15 |
-
|
| 16 |
-
## What Makes This Environment Non-Trivial
|
| 17 |
-
|
| 18 |
-
- Large decoy inventory: many resources are intentionally irrelevant.
|
| 19 |
-
- Ambiguous telemetry: medium/hard logs expose failing IPs, not only direct IDs.
|
| 20 |
-
- Action cost: every step incurs a small penalty, so efficient plans outperform exploratory spam.
|
| 21 |
-
- Hard-mode drift: delayed remediation causes additional system degradation.
|
| 22 |
-
|
| 23 |
-
## Command Reference
|
| 24 |
-
|
| 25 |
-
### list_resources
|
| 26 |
-
|
| 27 |
-
- Purpose: enumerate all available entities.
|
| 28 |
-
- Use first in most runs.
|
| 29 |
-
- Required fields: command only.
|
| 30 |
-
|
| 31 |
-
### describe_resource
|
| 32 |
-
|
| 33 |
-
- Purpose: inspect one resource configuration or current state.
|
| 34 |
-
- Required fields: resource_id.
|
| 35 |
-
|
| 36 |
-
### view_logs
|
| 37 |
-
|
| 38 |
-
- Purpose: read operational telemetry for one resource.
|
| 39 |
-
- Required fields: resource_id.
|
| 40 |
-
|
| 41 |
-
### query_metadata
|
| 42 |
-
|
| 43 |
-
- Purpose: resolve metadata such as ip_address -> resource_id.
|
| 44 |
-
- Required fields: parameters.ip_address.
|
| 45 |
-
- Typical medium/hard bridge action before mutation.
|
| 46 |
-
|
| 47 |
-
### update_security_group
|
| 48 |
|
| 49 |
-
|
| 50 |
-
- Required fields: resource_id, parameters.port, parameters.action.
|
| 51 |
-
- Valid action values: allow, deny.
|
| 52 |
|
| 53 |
-
##
|
| 54 |
|
| 55 |
-
-
|
| 56 |
-
|
| 57 |
-
- Use only after root-cause confirmation.
|
| 58 |
|
| 59 |
-
|
|
|
|
| 60 |
|
| 61 |
-
-
|
| 62 |
-
|
| 63 |
|
| 64 |
-
|
|
|
|
| 65 |
|
| 66 |
-
``
|
| 67 |
-
|
| 68 |
-
```
|
| 69 |
|
| 70 |
-
``
|
| 71 |
-
|
| 72 |
-
```
|
| 73 |
|
| 74 |
-
``
|
| 75 |
-
|
| 76 |
-
```
|
| 77 |
-
|
| 78 |
-
```json
|
| 79 |
-
{"command":"update_security_group","resource_id":"sg-db","parameters":{"port":5432,"action":"allow"}}
|
| 80 |
-
```
|
| 81 |
-
|
| 82 |
-
```json
|
| 83 |
-
{"command":"restart_service","resource_id":"i-web2"}
|
| 84 |
-
```
|
| 85 |
|
| 86 |
## Task Playbooks
|
| 87 |
|
| 88 |
-
### Easy
|
| 89 |
|
| 90 |
Objective:
|
| 91 |
-
|
| 92 |
-
- Open port 80 on sg-web.
|
| 93 |
|
| 94 |
Strong path:
|
|
|
|
|
|
|
| 95 |
|
| 96 |
-
|
| 97 |
-
2
|
| 98 |
-
|
| 99 |
-
Optional safer path (slightly lower efficiency):
|
| 100 |
|
| 101 |
-
|
| 102 |
-
2. describe_resource(sg-web)
|
| 103 |
-
3. update_security_group(sg-web, port=80, action=allow)
|
| 104 |
-
|
| 105 |
-
### Medium: API to DB Connectivity
|
| 106 |
|
| 107 |
Objective:
|
| 108 |
-
|
| 109 |
-
- Restore DB access by opening 5432 on sg-db, but only after diagnosis.
|
| 110 |
|
| 111 |
Strong path:
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
|
| 115 |
-
|
| 116 |
-
4. update_security_group(sg-db, port=5432, action=allow)
|
| 117 |
|
| 118 |
Guardrail:
|
|
|
|
| 119 |
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
### Hard: Upstream Failure Under Drift
|
| 123 |
|
| 124 |
Objective:
|
| 125 |
-
|
| 126 |
-
- Find failing upstream from lb-main logs, resolve target identity, inspect, then restart i-web2.
|
| 127 |
|
| 128 |
Strong path:
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
5. restart_service(i-web2)
|
| 135 |
|
| 136 |
Guardrails:
|
|
|
|
|
|
|
|
|
|
| 137 |
|
| 138 |
-
|
| 139 |
-
- Restarting i-web2 before investigation is penalized and blocked from resolution.
|
| 140 |
-
- If unresolved after step 8, lb-external also fails (cascading failure).
|
| 141 |
-
|
| 142 |
-
## Rewards, Termination, And Score Semantics
|
| 143 |
|
| 144 |
-
|
| 145 |
-
- Discovery and correct remediation provide shaped positive rewards.
|
| 146 |
-
- Unsafe or premature actions apply penalties.
|
| 147 |
-
- Episode terminates on resolution or max step limit.
|
| 148 |
-
- Submission score is clamped to a strict open interval (0,1).
|
| 149 |
-
- success reflects actual incident resolution state.
|
| 150 |
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
-
|
| 154 |
-
-
|
| 155 |
-
-
|
| 156 |
-
-
|
| 157 |
-
- done: episode termination flag
|
| 158 |
-
- metadata: task, step_count, resolved, achievements, action_cost
|
| 159 |
|
| 160 |
## Submission Contract Reminder
|
| 161 |
|
| 162 |
-
inference.py must:
|
| 163 |
|
| 164 |
- use OpenAI client
|
| 165 |
-
- read API_BASE_URL, MODEL_NAME, HF_TOKEN
|
| 166 |
-
- emit strict stdout
|
| 167 |
-
|
| 168 |
-
Practical strategy:
|
| 169 |
-
|
| 170 |
-
- Inspect first, mutate second.
|
| 171 |
-
- Use query_metadata whenever logs give only IP clues.
|
| 172 |
-
- Prefer minimal actions for higher score.
|
|
|
|
| 1 |
+
# Cloud DevOps RLEnv
|
| 2 |
|
| 3 |
+
Use this page as the operator playbook in the /web interface.
|
| 4 |
|
| 5 |
+
## Core Loop
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
|
| 7 |
1. Triage quickly.
|
| 8 |
+
2. Build evidence from logs and metadata.
|
| 9 |
+
3. Apply the minimum safe remediation.
|
| 10 |
+
4. Verify resolution and stop.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
|
| 12 |
+
Why this matters: each action has a cost (`-0.01`), so shorter correct trajectories score higher.
|
|
|
|
|
|
|
| 13 |
|
| 14 |
+
## Command Cheat Sheet
|
| 15 |
|
| 16 |
+
- `list_resources`
|
| 17 |
+
: Start here to map real targets vs decoys.
|
|
|
|
| 18 |
|
| 19 |
+
- `describe_resource(resource_id)`
|
| 20 |
+
: Inspect config/state details.
|
| 21 |
|
| 22 |
+
- `view_logs(resource_id)`
|
| 23 |
+
: Primary source of root-cause evidence.
|
| 24 |
|
| 25 |
+
- `query_metadata(parameters={"ip_address": "..."})`
|
| 26 |
+
: Resolve IP-only clues to resource IDs (mandatory multi-hop step in medium/hard).
|
| 27 |
|
| 28 |
+
- `update_security_group(resource_id, parameters)`
|
| 29 |
+
: Requires `parameters.port` and `parameters.action` where action is `allow` or `deny`.
|
|
|
|
| 30 |
|
| 31 |
+
- `restart_service(resource_id)`
|
| 32 |
+
: Use only after causal evidence confirms target.
|
|
|
|
| 33 |
|
| 34 |
+
- `submit_solution`
|
| 35 |
+
: Useful for explicit closure checks; unresolved hard submissions are penalized and continue.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
|
| 37 |
## Task Playbooks
|
| 38 |
|
| 39 |
+
### Easy
|
| 40 |
|
| 41 |
Objective:
|
| 42 |
+
- Restore web access by allowing port `80` on `sg-web`.
|
|
|
|
| 43 |
|
| 44 |
Strong path:
|
| 45 |
+
1. `list_resources`
|
| 46 |
+
2. `update_security_group("sg-web", {"port": 80, "action": "allow"})`
|
| 47 |
|
| 48 |
+
Typical outcome:
|
| 49 |
+
- 2 steps, score around `0.78`
|
|
|
|
|
|
|
| 50 |
|
| 51 |
+
### Medium
|
|
|
|
|
|
|
|
|
|
|
|
|
| 52 |
|
| 53 |
Objective:
|
| 54 |
+
- Restore API to DB connectivity by allowing port `5432` on `sg-db`.
|
|
|
|
| 55 |
|
| 56 |
Strong path:
|
| 57 |
+
1. `list_resources`
|
| 58 |
+
2. `view_logs("i-api")`
|
| 59 |
+
3. `query_metadata({"ip_address": "10.0.4.5"})`
|
| 60 |
+
4. `update_security_group("sg-db", {"port": 5432, "action": "allow"})`
|
|
|
|
| 61 |
|
| 62 |
Guardrail:
|
| 63 |
+
- Applying SG changes before logs + metadata lookup is penalized and does not resolve the incident.
|
| 64 |
|
| 65 |
+
### Hard
|
|
|
|
|
|
|
| 66 |
|
| 67 |
Objective:
|
| 68 |
+
- Recover checkout flow by identifying the failing upstream and restarting `i-web2` safely.
|
|
|
|
| 69 |
|
| 70 |
Strong path:
|
| 71 |
+
1. `list_resources`
|
| 72 |
+
2. `view_logs("lb-main")`
|
| 73 |
+
3. `query_metadata({"ip_address": "10.0.8.22"})`
|
| 74 |
+
4. `describe_resource("i-web2")` or `view_logs("i-web2")`
|
| 75 |
+
5. `restart_service("i-web2")`
|
|
|
|
| 76 |
|
| 77 |
Guardrails:
|
| 78 |
+
- Restarting `i-web1` is penalized.
|
| 79 |
+
- Restarting `i-web2` without investigation is penalized.
|
| 80 |
+
- If unresolved after step 8, `lb-external` also fails (cascading failure).
|
| 81 |
|
| 82 |
+
## Reading Metadata
|
|
|
|
|
|
|
|
|
|
|
|
|
| 83 |
|
| 84 |
+
Watch these response fields each step:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 85 |
|
| 86 |
+
- `system_health_status`: `CRITICAL` / `DEGRADED` / `HEALTHY`
|
| 87 |
+
- `done`: episode ended or still running
|
| 88 |
+
- `reward`: immediate signal after action cost and shaping
|
| 89 |
+
- `metadata.resolved`: authoritative success flag
|
| 90 |
+
- `metadata.termination_reason`: why episode ended
|
| 91 |
+
- `metadata.reward_breakdown`: transparent reward events for grader/debug inspection
|
|
|
|
|
|
|
| 92 |
|
| 93 |
## Submission Contract Reminder
|
| 94 |
|
| 95 |
+
`inference.py` must:
|
| 96 |
|
| 97 |
- use OpenAI client
|
| 98 |
+
- read `API_BASE_URL`, `MODEL_NAME`, `HF_TOKEN`
|
| 99 |
+
- emit strict stdout markers: `[START]`, `[STEP]`, `[END]`
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
scripts/reproducibility_check.py
ADDED
|
@@ -0,0 +1,125 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python
|
| 2 |
+
"""Deterministic reproducibility smoke test for Cloud DevOps RLEnv."""
|
| 3 |
+
|
| 4 |
+
from __future__ import annotations
|
| 5 |
+
|
| 6 |
+
import asyncio
|
| 7 |
+
import json
|
| 8 |
+
import sys
|
| 9 |
+
from pathlib import Path
|
| 10 |
+
from typing import Any
|
| 11 |
+
|
| 12 |
+
REPO_ROOT = Path(__file__).resolve().parent.parent
|
| 13 |
+
if str(REPO_ROOT) not in sys.path:
|
| 14 |
+
sys.path.insert(0, str(REPO_ROOT))
|
| 15 |
+
|
| 16 |
+
from env import CloudDevOpsEnv
|
| 17 |
+
from models import CloudAction
|
| 18 |
+
|
| 19 |
+
SCORE_MIN = 0.001
|
| 20 |
+
SCORE_MAX = 0.999
|
| 21 |
+
|
| 22 |
+
# Fixed trajectories that should always resolve the incidents.
|
| 23 |
+
POLICY_BY_TASK: dict[str, list[dict[str, Any]]] = {
|
| 24 |
+
"easy": [
|
| 25 |
+
{"command": "list_resources"},
|
| 26 |
+
{
|
| 27 |
+
"command": "update_security_group",
|
| 28 |
+
"resource_id": "sg-web",
|
| 29 |
+
"parameters": {"port": 80, "action": "allow"},
|
| 30 |
+
},
|
| 31 |
+
],
|
| 32 |
+
"medium": [
|
| 33 |
+
{"command": "list_resources"},
|
| 34 |
+
{"command": "view_logs", "resource_id": "i-api"},
|
| 35 |
+
{
|
| 36 |
+
"command": "query_metadata",
|
| 37 |
+
"parameters": {"ip_address": "10.0.4.5"},
|
| 38 |
+
},
|
| 39 |
+
{
|
| 40 |
+
"command": "update_security_group",
|
| 41 |
+
"resource_id": "sg-db",
|
| 42 |
+
"parameters": {"port": 5432, "action": "allow"},
|
| 43 |
+
},
|
| 44 |
+
],
|
| 45 |
+
"hard": [
|
| 46 |
+
{"command": "list_resources"},
|
| 47 |
+
{"command": "view_logs", "resource_id": "lb-main"},
|
| 48 |
+
{
|
| 49 |
+
"command": "query_metadata",
|
| 50 |
+
"parameters": {"ip_address": "10.0.8.22"},
|
| 51 |
+
},
|
| 52 |
+
{"command": "describe_resource", "resource_id": "i-web2"},
|
| 53 |
+
{"command": "restart_service", "resource_id": "i-web2"},
|
| 54 |
+
],
|
| 55 |
+
}
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
async def run_policy(task_name: str) -> dict[str, Any]:
|
| 59 |
+
env = CloudDevOpsEnv(task_name=task_name)
|
| 60 |
+
await env.reset()
|
| 61 |
+
|
| 62 |
+
trajectory: list[dict[str, Any]] = []
|
| 63 |
+
rewards: list[float] = []
|
| 64 |
+
last = None
|
| 65 |
+
|
| 66 |
+
try:
|
| 67 |
+
for index, raw_action in enumerate(POLICY_BY_TASK[task_name], start=1):
|
| 68 |
+
action = CloudAction(**raw_action)
|
| 69 |
+
result = await env.step(action)
|
| 70 |
+
rewards.append(float(result.reward))
|
| 71 |
+
trajectory.append(
|
| 72 |
+
{
|
| 73 |
+
"step": index,
|
| 74 |
+
"command": action.command,
|
| 75 |
+
"resource_id": action.resource_id,
|
| 76 |
+
"reward": round(float(result.reward), 4),
|
| 77 |
+
"done": bool(result.done),
|
| 78 |
+
"error": result.observation.error,
|
| 79 |
+
"status": result.observation.system_health_status,
|
| 80 |
+
"resolved": bool(result.info.get("resolved", False)),
|
| 81 |
+
}
|
| 82 |
+
)
|
| 83 |
+
last = result
|
| 84 |
+
if result.done:
|
| 85 |
+
break
|
| 86 |
+
|
| 87 |
+
if last is None:
|
| 88 |
+
raise RuntimeError(f"Task {task_name} produced no steps")
|
| 89 |
+
|
| 90 |
+
score = max(SCORE_MIN, min(sum(rewards), SCORE_MAX))
|
| 91 |
+
return {
|
| 92 |
+
"task": task_name,
|
| 93 |
+
"resolved": bool(last.info.get("resolved", False)),
|
| 94 |
+
"steps": len(trajectory),
|
| 95 |
+
"score": round(score, 3),
|
| 96 |
+
"trajectory": trajectory,
|
| 97 |
+
}
|
| 98 |
+
finally:
|
| 99 |
+
await env.close()
|
| 100 |
+
|
| 101 |
+
|
| 102 |
+
async def main() -> None:
|
| 103 |
+
summary: dict[str, Any] = {}
|
| 104 |
+
|
| 105 |
+
for task_name in ("easy", "medium", "hard"):
|
| 106 |
+
run_1 = await run_policy(task_name)
|
| 107 |
+
run_2 = await run_policy(task_name)
|
| 108 |
+
|
| 109 |
+
if run_1["trajectory"] != run_2["trajectory"]:
|
| 110 |
+
raise SystemExit(f"Determinism check failed for task={task_name}: trajectories differ")
|
| 111 |
+
if not run_1["resolved"]:
|
| 112 |
+
raise SystemExit(f"Policy failed to resolve task={task_name}")
|
| 113 |
+
|
| 114 |
+
summary[task_name] = {
|
| 115 |
+
"steps": run_1["steps"],
|
| 116 |
+
"score": run_1["score"],
|
| 117 |
+
"resolved": run_1["resolved"],
|
| 118 |
+
}
|
| 119 |
+
|
| 120 |
+
print("Deterministic reproducibility check passed")
|
| 121 |
+
print(json.dumps(summary, indent=2, sort_keys=True))
|
| 122 |
+
|
| 123 |
+
|
| 124 |
+
if __name__ == "__main__":
|
| 125 |
+
asyncio.run(main())
|
server/cloud_devops_env_environment.py
CHANGED
|
@@ -217,6 +217,20 @@ class CloudDevopsEnvironment(Environment):
|
|
| 217 |
self._achievements.add(achievement)
|
| 218 |
return points
|
| 219 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 220 |
def reset(self) -> CloudObservation: # type: ignore[override]
|
| 221 |
"""Reset the environment to the initial state for the selected task."""
|
| 222 |
self._achievements.clear()
|
|
@@ -242,6 +256,11 @@ class CloudDevopsEnvironment(Environment):
|
|
| 242 |
"resolved": False,
|
| 243 |
"task": self.task_name,
|
| 244 |
"total_resources": len(self._state_data.resources),
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 245 |
},
|
| 246 |
echoed_message="Cloud Devops Env environment ready!",
|
| 247 |
message_length=0,
|
|
@@ -257,9 +276,20 @@ class CloudDevopsEnvironment(Environment):
|
|
| 257 |
|
| 258 |
state.step_count += 1
|
| 259 |
reward = -self.ACTION_COST
|
|
|
|
|
|
|
|
|
|
| 260 |
done = False
|
| 261 |
output = ""
|
| 262 |
error = None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 263 |
|
| 264 |
try:
|
| 265 |
if action.command == "list_resources":
|
|
@@ -276,11 +306,14 @@ class CloudDevopsEnvironment(Environment):
|
|
| 276 |
output = str(state.resources[action.resource_id])
|
| 277 |
|
| 278 |
if self.task_name == "easy" and action.resource_id == "sg-web":
|
| 279 |
-
|
| 280 |
elif self.task_name == "medium" and action.resource_id == "sg-db":
|
| 281 |
-
|
| 282 |
elif self.task_name == "hard" and action.resource_id == "i-web2":
|
| 283 |
-
|
|
|
|
|
|
|
|
|
|
| 284 |
|
| 285 |
elif action.command == "view_logs":
|
| 286 |
if not action.resource_id:
|
|
@@ -293,11 +326,14 @@ class CloudDevopsEnvironment(Environment):
|
|
| 293 |
output = str(res.get("logs", "No logs available for this resource."))
|
| 294 |
|
| 295 |
if self.task_name == "medium" and action.resource_id == "i-api":
|
| 296 |
-
|
| 297 |
elif self.task_name == "hard" and action.resource_id == "lb-main":
|
| 298 |
-
|
| 299 |
elif self.task_name == "hard" and action.resource_id == "i-web2":
|
| 300 |
-
|
|
|
|
|
|
|
|
|
|
| 301 |
|
| 302 |
elif action.command == "query_metadata":
|
| 303 |
ip_address = None
|
|
@@ -314,9 +350,15 @@ class CloudDevopsEnvironment(Environment):
|
|
| 314 |
|
| 315 |
output = f"Metadata lookup: ip_address={ip_address} resource_id={resource_id}"
|
| 316 |
if self.task_name == "medium" and str(ip_address) == "10.0.4.5":
|
| 317 |
-
|
|
|
|
|
|
|
|
|
|
| 318 |
elif self.task_name == "hard" and str(ip_address) == "10.0.8.22":
|
| 319 |
-
|
|
|
|
|
|
|
|
|
|
| 320 |
|
| 321 |
elif action.command == "update_security_group":
|
| 322 |
if not action.resource_id:
|
|
@@ -350,8 +392,9 @@ class CloudDevopsEnvironment(Environment):
|
|
| 350 |
and rule_action == "allow"
|
| 351 |
):
|
| 352 |
state.is_resolved = True
|
| 353 |
-
|
| 354 |
done = True
|
|
|
|
| 355 |
output += "\nSUCCESS: Web server is now accessible!"
|
| 356 |
elif (
|
| 357 |
self.task_name == "medium"
|
|
@@ -365,17 +408,18 @@ class CloudDevopsEnvironment(Environment):
|
|
| 365 |
)
|
| 366 |
if investigated:
|
| 367 |
state.is_resolved = True
|
| 368 |
-
|
| 369 |
done = True
|
|
|
|
| 370 |
output += "\nSUCCESS: Database connection restored!"
|
| 371 |
else:
|
| 372 |
-
|
| 373 |
output += (
|
| 374 |
"\nWARNING: Change applied without incident triage. "
|
| 375 |
"Inspect API logs and resolve DB IP via query_metadata before closing the incident."
|
| 376 |
)
|
| 377 |
elif rule_action == "deny":
|
| 378 |
-
|
| 379 |
output += "\nWARNING: Deny rule applied during outage remediation."
|
| 380 |
|
| 381 |
elif action.command == "restart_service":
|
|
@@ -399,17 +443,18 @@ class CloudDevopsEnvironment(Environment):
|
|
| 399 |
"logs"
|
| 400 |
] = "INFO: Restart successful. Memory cleared."
|
| 401 |
state.is_resolved = True
|
| 402 |
-
|
| 403 |
done = True
|
|
|
|
| 404 |
output += "\nSUCCESS: OutOfMemory loop broken. System stable."
|
| 405 |
else:
|
| 406 |
-
|
| 407 |
output += (
|
| 408 |
"\nWARNING: Restart denied by change policy. "
|
| 409 |
"Find failing upstream IP from lb-main, resolve it with query_metadata, and inspect i-web2 first."
|
| 410 |
)
|
| 411 |
elif action.resource_id == "i-web1":
|
| 412 |
-
|
| 413 |
output += (
|
| 414 |
"\nWARNING: You restarted a healthy production server! "
|
| 415 |
"Users dropped."
|
|
@@ -418,18 +463,20 @@ class CloudDevopsEnvironment(Environment):
|
|
| 418 |
elif action.command == "submit_solution":
|
| 419 |
if state.is_resolved:
|
| 420 |
done = True
|
|
|
|
| 421 |
output = "Solution verified. System is HEALTHY."
|
| 422 |
else:
|
| 423 |
if self.task_name == "hard":
|
| 424 |
# In hard mode, unresolved submission should not abort the run.
|
| 425 |
done = False
|
| 426 |
-
|
| 427 |
output = (
|
| 428 |
"Solution incorrect. Incident is still CRITICAL. "
|
| 429 |
"Continue triage and remediation before submitting."
|
| 430 |
)
|
| 431 |
else:
|
| 432 |
done = True
|
|
|
|
| 433 |
output = "Solution incorrect. System is still CRITICAL."
|
| 434 |
|
| 435 |
else:
|
|
@@ -440,16 +487,26 @@ class CloudDevopsEnvironment(Environment):
|
|
| 440 |
output = f"Command Failed: {error}"
|
| 441 |
|
| 442 |
cascade_penalty, cascade_msg = self._apply_cascading_failure()
|
| 443 |
-
|
| 444 |
if cascade_msg:
|
| 445 |
output = f"{output}{cascade_msg}" if output else cascade_msg.strip()
|
| 446 |
|
| 447 |
if state.step_count >= self.MAX_STEPS and not done:
|
| 448 |
done = True
|
|
|
|
| 449 |
timeout_suffix = "\nTIMEOUT: Max steps reached."
|
| 450 |
output = f"{output}{timeout_suffix}" if output else timeout_suffix.strip()
|
| 451 |
|
|
|
|
| 452 |
reward = max(-1.0, min(1.0, reward))
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 453 |
lb_external = state.resources.get("lb-external", {})
|
| 454 |
if state.is_resolved:
|
| 455 |
status = "HEALTHY"
|
|
@@ -464,6 +521,11 @@ class CloudDevopsEnvironment(Environment):
|
|
| 464 |
"achievements": sorted(self._achievements),
|
| 465 |
"total_resources": len(state.resources),
|
| 466 |
"action_cost": self.ACTION_COST,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 467 |
}
|
| 468 |
|
| 469 |
return CloudObservation(
|
|
|
|
| 217 |
self._achievements.add(achievement)
|
| 218 |
return points
|
| 219 |
|
| 220 |
+
def _task_objective(self) -> str:
|
| 221 |
+
objectives = {
|
| 222 |
+
"easy": "Restore web access by allowing port 80 on sg-web.",
|
| 223 |
+
"medium": (
|
| 224 |
+
"Restore API to DB connectivity by reading i-api logs, resolving DB IP via "
|
| 225 |
+
"query_metadata, then allowing port 5432 on sg-db."
|
| 226 |
+
),
|
| 227 |
+
"hard": (
|
| 228 |
+
"Recover checkout path by tracing lb-main upstream IP, resolving it with "
|
| 229 |
+
"query_metadata, inspecting i-web2, and restarting i-web2 safely."
|
| 230 |
+
),
|
| 231 |
+
}
|
| 232 |
+
return objectives[self.task_name]
|
| 233 |
+
|
| 234 |
def reset(self) -> CloudObservation: # type: ignore[override]
|
| 235 |
"""Reset the environment to the initial state for the selected task."""
|
| 236 |
self._achievements.clear()
|
|
|
|
| 256 |
"resolved": False,
|
| 257 |
"task": self.task_name,
|
| 258 |
"total_resources": len(self._state_data.resources),
|
| 259 |
+
"objective": self._task_objective(),
|
| 260 |
+
"deterministic": True,
|
| 261 |
+
"max_steps": self.MAX_STEPS,
|
| 262 |
+
"action_cost": self.ACTION_COST,
|
| 263 |
+
"hard_cascade_trigger_step": 8,
|
| 264 |
},
|
| 265 |
echoed_message="Cloud Devops Env environment ready!",
|
| 266 |
message_length=0,
|
|
|
|
| 276 |
|
| 277 |
state.step_count += 1
|
| 278 |
reward = -self.ACTION_COST
|
| 279 |
+
reward_breakdown: list[dict[str, object]] = [
|
| 280 |
+
{"event": "action_cost", "delta": -self.ACTION_COST}
|
| 281 |
+
]
|
| 282 |
done = False
|
| 283 |
output = ""
|
| 284 |
error = None
|
| 285 |
+
termination_reason = "in_progress"
|
| 286 |
+
|
| 287 |
+
def add_reward(delta: float, event: str) -> None:
|
| 288 |
+
nonlocal reward
|
| 289 |
+
if abs(delta) < 1e-12:
|
| 290 |
+
return
|
| 291 |
+
reward += delta
|
| 292 |
+
reward_breakdown.append({"event": event, "delta": round(float(delta), 4)})
|
| 293 |
|
| 294 |
try:
|
| 295 |
if action.command == "list_resources":
|
|
|
|
| 306 |
output = str(state.resources[action.resource_id])
|
| 307 |
|
| 308 |
if self.task_name == "easy" and action.resource_id == "sg-web":
|
| 309 |
+
add_reward(self._reward_once("read_sg", 0.2), "inspect_web_sg")
|
| 310 |
elif self.task_name == "medium" and action.resource_id == "sg-db":
|
| 311 |
+
add_reward(self._reward_once("read_sg", 0.2), "inspect_db_sg")
|
| 312 |
elif self.task_name == "hard" and action.resource_id == "i-web2":
|
| 313 |
+
add_reward(
|
| 314 |
+
self._reward_once("inspect_target", 0.2),
|
| 315 |
+
"inspect_target_instance",
|
| 316 |
+
)
|
| 317 |
|
| 318 |
elif action.command == "view_logs":
|
| 319 |
if not action.resource_id:
|
|
|
|
| 326 |
output = str(res.get("logs", "No logs available for this resource."))
|
| 327 |
|
| 328 |
if self.task_name == "medium" and action.resource_id == "i-api":
|
| 329 |
+
add_reward(self._reward_once("read_logs", 0.2), "inspect_api_logs")
|
| 330 |
elif self.task_name == "hard" and action.resource_id == "lb-main":
|
| 331 |
+
add_reward(self._reward_once("inspect_lb", 0.2), "inspect_lb_logs")
|
| 332 |
elif self.task_name == "hard" and action.resource_id == "i-web2":
|
| 333 |
+
add_reward(
|
| 334 |
+
self._reward_once("inspect_target", 0.2),
|
| 335 |
+
"inspect_target_logs",
|
| 336 |
+
)
|
| 337 |
|
| 338 |
elif action.command == "query_metadata":
|
| 339 |
ip_address = None
|
|
|
|
| 350 |
|
| 351 |
output = f"Metadata lookup: ip_address={ip_address} resource_id={resource_id}"
|
| 352 |
if self.task_name == "medium" and str(ip_address) == "10.0.4.5":
|
| 353 |
+
add_reward(
|
| 354 |
+
self._reward_once("lookup_db_target", 0.2),
|
| 355 |
+
"resolve_db_ip_dependency",
|
| 356 |
+
)
|
| 357 |
elif self.task_name == "hard" and str(ip_address) == "10.0.8.22":
|
| 358 |
+
add_reward(
|
| 359 |
+
self._reward_once("lookup_upstream_target", 0.2),
|
| 360 |
+
"resolve_upstream_ip_dependency",
|
| 361 |
+
)
|
| 362 |
|
| 363 |
elif action.command == "update_security_group":
|
| 364 |
if not action.resource_id:
|
|
|
|
| 392 |
and rule_action == "allow"
|
| 393 |
):
|
| 394 |
state.is_resolved = True
|
| 395 |
+
add_reward(0.8, "resolve_easy_web_ingress")
|
| 396 |
done = True
|
| 397 |
+
termination_reason = "resolved_easy"
|
| 398 |
output += "\nSUCCESS: Web server is now accessible!"
|
| 399 |
elif (
|
| 400 |
self.task_name == "medium"
|
|
|
|
| 408 |
)
|
| 409 |
if investigated:
|
| 410 |
state.is_resolved = True
|
| 411 |
+
add_reward(0.6, "resolve_medium_db_connectivity")
|
| 412 |
done = True
|
| 413 |
+
termination_reason = "resolved_medium"
|
| 414 |
output += "\nSUCCESS: Database connection restored!"
|
| 415 |
else:
|
| 416 |
+
add_reward(-0.1, "unsafe_change_without_triage")
|
| 417 |
output += (
|
| 418 |
"\nWARNING: Change applied without incident triage. "
|
| 419 |
"Inspect API logs and resolve DB IP via query_metadata before closing the incident."
|
| 420 |
)
|
| 421 |
elif rule_action == "deny":
|
| 422 |
+
add_reward(-0.1, "deny_rule_during_incident")
|
| 423 |
output += "\nWARNING: Deny rule applied during outage remediation."
|
| 424 |
|
| 425 |
elif action.command == "restart_service":
|
|
|
|
| 443 |
"logs"
|
| 444 |
] = "INFO: Restart successful. Memory cleared."
|
| 445 |
state.is_resolved = True
|
| 446 |
+
add_reward(0.8, "resolve_hard_upstream_recovery")
|
| 447 |
done = True
|
| 448 |
+
termination_reason = "resolved_hard"
|
| 449 |
output += "\nSUCCESS: OutOfMemory loop broken. System stable."
|
| 450 |
else:
|
| 451 |
+
add_reward(-0.1, "restart_without_root_cause")
|
| 452 |
output += (
|
| 453 |
"\nWARNING: Restart denied by change policy. "
|
| 454 |
"Find failing upstream IP from lb-main, resolve it with query_metadata, and inspect i-web2 first."
|
| 455 |
)
|
| 456 |
elif action.resource_id == "i-web1":
|
| 457 |
+
add_reward(-0.2, "restart_healthy_node")
|
| 458 |
output += (
|
| 459 |
"\nWARNING: You restarted a healthy production server! "
|
| 460 |
"Users dropped."
|
|
|
|
| 463 |
elif action.command == "submit_solution":
|
| 464 |
if state.is_resolved:
|
| 465 |
done = True
|
| 466 |
+
termination_reason = "resolved_submit_solution"
|
| 467 |
output = "Solution verified. System is HEALTHY."
|
| 468 |
else:
|
| 469 |
if self.task_name == "hard":
|
| 470 |
# In hard mode, unresolved submission should not abort the run.
|
| 471 |
done = False
|
| 472 |
+
add_reward(-0.1, "premature_submit_hard")
|
| 473 |
output = (
|
| 474 |
"Solution incorrect. Incident is still CRITICAL. "
|
| 475 |
"Continue triage and remediation before submitting."
|
| 476 |
)
|
| 477 |
else:
|
| 478 |
done = True
|
| 479 |
+
termination_reason = "incorrect_submit"
|
| 480 |
output = "Solution incorrect. System is still CRITICAL."
|
| 481 |
|
| 482 |
else:
|
|
|
|
| 487 |
output = f"Command Failed: {error}"
|
| 488 |
|
| 489 |
cascade_penalty, cascade_msg = self._apply_cascading_failure()
|
| 490 |
+
add_reward(cascade_penalty, "cascading_failure_penalty")
|
| 491 |
if cascade_msg:
|
| 492 |
output = f"{output}{cascade_msg}" if output else cascade_msg.strip()
|
| 493 |
|
| 494 |
if state.step_count >= self.MAX_STEPS and not done:
|
| 495 |
done = True
|
| 496 |
+
termination_reason = "max_steps_timeout"
|
| 497 |
timeout_suffix = "\nTIMEOUT: Max steps reached."
|
| 498 |
output = f"{output}{timeout_suffix}" if output else timeout_suffix.strip()
|
| 499 |
|
| 500 |
+
raw_reward = reward
|
| 501 |
reward = max(-1.0, min(1.0, reward))
|
| 502 |
+
if reward != raw_reward:
|
| 503 |
+
reward_breakdown.append(
|
| 504 |
+
{
|
| 505 |
+
"event": "reward_clip",
|
| 506 |
+
"delta": round(float(reward - raw_reward), 4),
|
| 507 |
+
}
|
| 508 |
+
)
|
| 509 |
+
|
| 510 |
lb_external = state.resources.get("lb-external", {})
|
| 511 |
if state.is_resolved:
|
| 512 |
status = "HEALTHY"
|
|
|
|
| 521 |
"achievements": sorted(self._achievements),
|
| 522 |
"total_resources": len(state.resources),
|
| 523 |
"action_cost": self.ACTION_COST,
|
| 524 |
+
"objective": self._task_objective(),
|
| 525 |
+
"deterministic": True,
|
| 526 |
+
"max_steps": self.MAX_STEPS,
|
| 527 |
+
"termination_reason": termination_reason if done else "in_progress",
|
| 528 |
+
"reward_breakdown": reward_breakdown,
|
| 529 |
}
|
| 530 |
|
| 531 |
return CloudObservation(
|