SidhaGarg commited on
Commit
f97f9a6
·
1 Parent(s): 1053889

Strengthen grading transparency and reproducibility evidence

Browse files
README.md CHANGED
@@ -11,158 +11,273 @@ tags:
11
 
12
  # Cloud DevOps RLEnv
13
 
14
- Cloud DevOps RLEnv is an OpenEnv benchmark for real incident-response workflows in cloud production systems.
15
 
16
- It is designed to evaluate whether an agent can triage noisy telemetry, identify root cause, apply a safe fix, and do all of this efficiently under cost and pressure.
17
 
18
  ## Judge-Aligned Snapshot
19
 
20
- This section maps the environment directly to the scoring rubric.
21
-
22
  | Parameter | Weight | How this environment addresses it |
23
  | --- | --- | --- |
24
- | Real-world utility | 30% | Simulates practical SRE outage response loops: triage logs, map failing dependency, remediate safely, verify health. |
25
- | Task and grader quality | 25% | Three deterministic tasks (easy/medium/hard), explicit objectives, reproducible reward logic, strict success/failure gates. |
26
- | Environment design | 20% | Typed actions and observations, deterministic reset, shaped rewards, action cost, clear episode boundaries and timeout. |
27
- | Code quality and spec compliance | 15% | OpenEnv-compatible layout, typed models, Dockerized runtime, strict inference logging contract, local validators included. |
28
- | Creativity and novelty | 10% | Multi-hop metadata dependency, cascading failure drift, high-decoy infrastructure search, efficiency-aware reward shaping. |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
29
 
30
- ## Real-World Utility
 
 
 
 
31
 
32
- The benchmark models incidents that closely mirror day-2 operations in production:
33
 
34
- - Security group misconfiguration blocks service traffic.
35
- - Service-to-database communication fails with telemetry-first diagnosis required.
36
- - Load balancer upstream failures require dependency mapping and targeted restart under time pressure.
 
 
 
 
 
 
 
 
 
 
 
 
37
 
38
- This is not a toy command simulator. The hard task requires reasoning over partial evidence and acting safely under drift.
39
 
40
- ## Why It Is Hard
 
 
41
 
42
- - Needle-in-a-haystack discovery: 20+ decoy instances and 20+ decoy security groups.
43
- - Ambiguous telemetry: logs expose symptoms and IPs, not always direct resource names.
44
- - Action-cost pressure: every action incurs a negative reward, penalizing brute-force behavior.
45
- - Multi-hop dependency: agent must use metadata resolution before remediation in medium/hard.
46
- - Cascading failures: unresolved hard incidents degrade additional components after step 8.
47
 
48
- ## Task Suite And Difficulty Progression
49
 
50
- | Task | Objective | Required reasoning depth | Typical strong trajectory |
51
  | --- | --- | --- | --- |
52
- | easy | Restore web access by opening port 80 on sg-web | Single-hop config diagnosis | list_resources -> update_security_group(sg-web, allow 80) |
53
- | medium | Restore DB connectivity by opening port 5432 on sg-db | Multi-hop: logs -> IP -> metadata lookup -> fix | list_resources -> view_logs(i-api) -> query_metadata(10.0.4.5) -> update_security_group(sg-db, allow 5432) |
54
- | hard | Recover checkout path by fixing failing upstream i-web2 | Multi-hop + pressure: logs -> IP -> metadata -> inspect -> restart | list_resources -> view_logs(lb-main) -> query_metadata(10.0.8.22) -> describe/view i-web2 -> restart_service(i-web2) |
55
 
56
- ## Grader Quality And Determinism
 
 
 
 
 
 
 
57
 
58
- - Deterministic state: no RNG in task initialization.
59
- - Deterministic transitions: action handling is rule-based and reproducible.
60
- - Deterministic rewards: shaped by explicit achievement checkpoints and safety penalties.
61
- - Deterministic success: inferred from resolved incident state, not ad hoc heuristics.
62
 
63
- Score behavior:
64
 
65
- - Step reward is clipped to [-1.0, 1.0].
66
- - Per-task inference score is clamped to strict open interval (0, 1) for validator compatibility.
67
- - Lower-step successful solutions naturally score higher due to per-step action cost.
 
 
 
 
 
68
 
69
- ## Environment Design Details
70
 
71
- - Runtime health states: CRITICAL, DEGRADED, HEALTHY.
72
- - Episode limit: 20 environment steps.
73
- - Hard-task drift: if unresolved after step 8, lb-external is marked DOWN.
74
- - Action cost: -0.01 on every step to encourage efficient plans.
 
 
75
 
76
- ### Action Space (CloudAction)
77
 
78
- | Field | Type | Required | Notes |
79
- | --- | --- | --- | --- |
80
- | command | enum | yes | list_resources, describe_resource, view_logs, query_metadata, update_security_group, restart_service, submit_solution |
81
- | resource_id | string | conditional | Required for most commands except list_resources and query_metadata |
82
- | parameters | object | conditional | Required for query_metadata.ip_address and update_security_group.port/action |
 
83
 
84
- Security-group mutation semantics:
 
85
 
86
- - update_security_group requires both port and action.
87
- - action must be allow or deny.
88
- - Invalid actions are rejected and do not mutate state.
89
 
90
- ### Observation and State Models
91
 
92
- Observation (CloudObservation): output, error, system_health_status, reward, done, metadata.
 
93
 
94
- State (CloudState): task_difficulty, resources, step_count, is_resolved.
 
95
 
96
- ## Inference Contract And Compliance
 
 
 
97
 
98
- Mandatory environment variables:
 
 
99
 
100
- - API_BASE_URL
101
- - MODEL_NAME
102
- - HF_TOKEN
103
 
104
- Mandatory inference requirements:
 
105
 
106
- - inference.py is at repository root.
107
- - Uses OpenAI client for all LLM calls.
108
- - Emits strict stdout contract only:
109
 
110
- ```text
111
- [START] task=<task_name> env=<benchmark> model=<model_name>
112
- [STEP] step=<n> action=<action_json_or_str> reward=<0.00> done=<true|false> error=<msg|null>
113
- [END] success=<true|false> steps=<n> score=<score> rewards=<r1,r2,...,rn>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
114
  ```
115
 
116
- ## Reference Baseline Behavior
117
 
118
- Recent deterministic policy-style run with current environment dynamics:
 
119
 
120
- | Task | Steps | Score | Outcome |
121
- | --- | --- | --- | --- |
122
- | easy | 2 | 0.780 | success=true |
123
- | medium | 4 | 0.960 | success=true |
124
- | hard | 5 | 0.999 | success=true |
125
 
126
- ## Project Layout And Spec Files
 
 
 
 
 
127
 
128
- Required OpenEnv files are present:
 
 
 
 
129
 
130
- - openenv.yaml
131
- - env.py
132
- - models.py
133
- - inference.py
134
- - server/app.py
135
- - server/cloud_devops_env_environment.py
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
136
 
137
- ## Validation And Reproducibility
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
138
 
139
  From repository root:
140
 
141
  ```bash
142
- # OpenEnv manifest and schema checks
143
  ..\\.venv\\Scripts\\openenv validate
144
 
145
- # Full submission checks including inference contract
146
- bash scripts/pre_submit_validate.sh --ping-url https://<your-space>.hf.space
 
 
 
 
 
 
 
147
 
148
- # Official 3-step baseline validator
149
- bash scripts/validate-submission.sh https://<your-space>.hf.space .
150
 
151
- # Docker smoke test
152
- docker build -t cloud-devops-env:latest -f Dockerfile .
153
- docker run --rm -p 8000:8000 cloud-devops-env:latest
154
  ```
155
 
156
- ## Hugging Face Space Deployment Checklist
157
 
158
- 1. Keep front matter intact, including openenv tag and app_port 8000.
159
- 2. Push latest main branch to the Space repository.
160
- 3. Configure API_BASE_URL, MODEL_NAME, HF_TOKEN in Space settings.
161
- 4. Wait for build completion and verify endpoints:
162
- - GET /health -> 200
163
- - POST /reset -> 200
164
- 5. Run inference.py once in Space logs to confirm strict START/STEP/END output.
 
 
 
165
 
166
  Reference:
167
-
168
  - https://huggingface.co/docs/hub/spaces-config-reference
 
11
 
12
  # Cloud DevOps RLEnv
13
 
14
+ Cloud DevOps RLEnv is an OpenEnv-compatible cloud incident-response benchmark designed for agentic SRE and DevOps workflows.
15
 
16
+ This environment rewards correct diagnosis and safe remediation, not blind action execution. It is deterministic, reproducible, and optimized for hackathon evaluation.
17
 
18
  ## Judge-Aligned Snapshot
19
 
 
 
20
  | Parameter | Weight | How this environment addresses it |
21
  | --- | --- | --- |
22
+ | Real-world utility | 30% | Models practical SRE outage response loops: telemetry triage, dependency mapping, and safe remediation. |
23
+ | Task & grader quality | 25% | Three deterministic tasks with explicit objectives, strict success gates, and reproducible scoring behavior. |
24
+ | Environment design | 20% | Typed action/observation/state models, clean reset semantics, shaped rewards, action-cost efficiency pressure, clear boundaries. |
25
+ | Code quality & spec compliance | 15% | OpenEnv-compliant project layout, Dockerized runtime, strict inference output contract, validation scripts. |
26
+ | Creativity & novelty | 10% | Multi-hop metadata dependency, cascading failures, high-decoy search space, and safety-aware penalties. |
27
+
28
+ ## Why This Environment
29
+
30
+ Real incidents are multi-step and noisy. Good agents must:
31
+ - gather context before changing systems
32
+ - identify root cause from logs and topology
33
+ - apply minimal, correct fixes
34
+ - verify resolution
35
+
36
+ Cloud DevOps RLEnv simulates that behavior with realistic failure patterns, decoy resources, shaped rewards, and anti-shortcut guardrails.
37
+
38
+ ## Why It's Hard
39
+
40
+ This benchmark is intentionally designed to resist brute-force policies and reward disciplined SRE reasoning:
41
 
42
+ - Needle-in-a-haystack discovery: 20+ decoy compute nodes and 20+ decoy security groups increase search complexity.
43
+ - Ambiguous telemetry: noisy, raw operational logs surface symptoms (including IP-only clues) rather than direct root-cause labels.
44
+ - Action-penalty heuristics: every action has a small negative cost, so efficient remediation beats command spamming.
45
+ - Multi-hop dependency resolution: agents must map IP addresses to resource IDs via metadata lookup before applying fixes.
46
+ - System drift under pressure: in hard mode, delayed remediation triggers cascading failures that worsen observability and reward dynamics.
47
 
48
+ ## Environment Scope
49
 
50
+ - Domain: Cloud SRE / DevOps incident response
51
+ - Difficulty tiers: easy, medium, hard
52
+ - Max environment steps per episode: 20
53
+ - Runtime health states: CRITICAL, DEGRADED, HEALTHY
54
+ - Decoy resources: 20 backend instances + 20 backend security groups
55
+
56
+ ## OpenEnv Compliance
57
+
58
+ Core files:
59
+ - openenv.yaml
60
+ - env.py
61
+ - models.py
62
+ - inference.py
63
+ - server/app.py
64
+ - server/cloud_devops_env_environment.py
65
 
66
+ Validator command:
67
 
68
+ ```bash
69
+ ..\\.venv\\Scripts\\openenv validate
70
+ ```
71
 
72
+ ## Action Space
 
 
 
 
73
 
74
+ Model: CloudAction
75
 
76
+ | Field | Type | Required | Description |
77
  | --- | --- | --- | --- |
78
+ | command | enum | yes | One of: list_resources, describe_resource, view_logs, query_metadata, update_security_group, restart_service, submit_solution |
79
+ | resource_id | string | conditional | Required for most actions except list_resources and query_metadata |
80
+ | parameters | object | conditional | Used by mutating actions (for example, security-group updates) |
81
 
82
+ Action semantics:
83
+ - list_resources: Enumerates available resources including decoys.
84
+ - describe_resource: Returns structured details for one resource.
85
+ - view_logs: Returns logs for one resource.
86
+ - query_metadata: Resolves infrastructure metadata (for example, IP address to resource ID).
87
+ - update_security_group: Appends a rule (requires parameters.port and parameters.action where action is allow/deny).
88
+ - restart_service: Restarts one instance/service by ID.
89
+ - submit_solution: Declares the episode solved (or not solved).
90
 
91
+ ## Observation And State Space
 
 
 
92
 
93
+ Observation model: CloudObservation
94
 
95
+ | Field | Description |
96
+ | --- | --- |
97
+ | output | Main command output |
98
+ | error | Error string for failed commands |
99
+ | system_health_status | CRITICAL, DEGRADED, HEALTHY |
100
+ | done | Episode terminal flag |
101
+ | reward | Step reward |
102
+ | metadata | Diagnostics such as task, step_count, resolved, achievements |
103
 
104
+ Hidden state model: CloudState
105
 
106
+ | Field | Description |
107
+ | --- | --- |
108
+ | task_difficulty | easy, medium, hard |
109
+ | resources | Full resource graph including logs/rules |
110
+ | step_count | Current step counter |
111
+ | is_resolved | Whether root cause has been fixed |
112
 
113
+ ## Reward Design
114
 
115
+ Reward shaping is sparse-but-guided:
116
+ - discovery rewards for correct investigative steps
117
+ - larger terminal rewards for correct remediation
118
+ - penalties for unsafe or premature operations
119
+ - fixed action cost per step (efficiency pressure)
120
+ - timeout terminal condition after max steps
121
 
122
+ Per-step reward is clipped to [-1.0, 1.0].
123
+ Inference task score is adjusted to remain strictly within (0.0, 1.0) for Phase-2 validator compatibility.
124
 
125
+ ## Detailed Task Playbooks
 
 
126
 
127
+ ### Easy Task
128
 
129
+ Incident:
130
+ - Web traffic blocked by security group.
131
 
132
+ Objective:
133
+ - Open port 80 on sg-web.
134
 
135
+ Typical successful sequence:
136
+ 1. list_resources
137
+ 2. describe_resource(sg-web) for context (+0.2)
138
+ 3. update_security_group(sg-web, port=80, action=allow) (+0.8, done)
139
 
140
+ Expected score:
141
+ - ~0.97 for full playbook with efficient triage
142
+ - ~0.79 if agent skips the optional read step
143
 
144
+ ### Medium Task
 
 
145
 
146
+ Incident:
147
+ - API cannot reach DB due to blocked port 5432.
148
 
149
+ Objective:
150
+ - Confirm root cause from logs, then open port 5432 on sg-db.
 
151
 
152
+ Typical successful sequence:
153
+ 1. list_resources
154
+ 2. view_logs(i-api) to identify DB timeout (+0.2)
155
+ 3. query_metadata(ip_address=10.0.4.5) to resolve DB target (+0.2)
156
+ 4. update_security_group(sg-db, port=5432, action=allow) (+0.6, done if logs and metadata lookup were completed)
157
+
158
+ Guardrail:
159
+ - Applying the SG change before log triage + metadata lookup gives a penalty (-0.1) and does not close the incident.
160
+
161
+ Expected score:
162
+ - ~0.97 with full investigative path (logs -> metadata lookup -> remediation)
163
+ - below ~0.90 when metadata dependency is skipped
164
+
165
+ ## Determinism And Grader Transparency
166
+
167
+ - Deterministic reset and transitions: no randomization is used in task generation or transition logic.
168
+ - Transparent grading signals: observation metadata includes achievements, resolution status, termination reason, and reward breakdown events.
169
+ - Reproducibility helper:
170
+
171
+ ```bash
172
+ ..\\.venv\\Scripts\\python scripts/reproducibility_check.py
173
  ```
174
 
175
+ ### Hard Task
176
 
177
+ Incident:
178
+ - Checkout path degraded due to upstream timeout to an IP-only target that must be resolved first.
179
 
180
+ Objective:
181
+ - Trace LB errors to the correct target, resolve resource identity via metadata, and restart i-web2 only after diagnosis.
 
 
 
182
 
183
+ Typical successful sequence:
184
+ 1. list_resources
185
+ 2. view_logs(lb-main) to identify failing upstream IP (+0.2)
186
+ 3. query_metadata(ip_address=<failing_ip>) to resolve target ID (+0.2)
187
+ 4. describe_resource(i-web2) or view_logs(i-web2) (+0.2)
188
+ 5. restart_service(i-web2) (+0.8, done when all investigation achievements exist)
189
 
190
+ Guardrails:
191
+ - Restarting i-web2 before investigation: penalty (-0.1), no resolution.
192
+ - Restarting healthy i-web1: penalty (-0.2).
193
+ - Premature submit_solution in hard mode: penalty (-0.1), episode continues.
194
+ - If unresolved after step 8 in hard mode, lb-external also fails (cascading failure), increasing pressure and noise.
195
 
196
+ Expected score:
197
+ - near 1.0 after score clamping for strong trajectories (can exceed 1.0 raw before clamp)
198
+
199
+ ## API Endpoints
200
+
201
+ Core runtime:
202
+ - GET /health
203
+ - POST /reset
204
+ - POST /step
205
+ - GET /state
206
+ - GET /schema
207
+ - WS /ws
208
+
209
+ Web UI runtime:
210
+ - GET /web
211
+ - POST /web/reset
212
+ - POST /web/step
213
+ - GET /web/state
214
+ - GET /web/metadata
215
+
216
+ ## Inference Contract
217
+
218
+ inference.py requirements:
219
+ - uses OpenAI client
220
+ - reads API_BASE_URL, MODEL_NAME, HF_TOKEN
221
+ - emits strict logs: [START], [STEP], [END]
222
 
223
+ Current defaults in code:
224
+ - MODEL_NAME default: google/gemma-4-26B-A4B-it
225
+ - MAX_STEPS (in inference loop): 15
226
+ - success flag is derived from environment resolution state
227
+
228
+ ## Baselines
229
+
230
+ ### Deterministic task baselines
231
+
232
+ | Task | Typical baseline score |
233
+ | --- | --- |
234
+ | easy | 0.78 |
235
+ | medium | 0.96 |
236
+ | hard | 0.999 |
237
+
238
+ ### LLM policy comparison
239
+
240
+ | Model | Easy | Medium | Summary |
241
+ | --- | --- | --- | --- |
242
+ | gemma-3-27b-it | 0.2 | 0.2 | Underperformed on this environment |
243
+ | gemma-4-31b-it | 1.0 | 1.0 | Perfect on both easy and medium |
244
+
245
+ ## Local Setup And Validation
246
 
247
  From repository root:
248
 
249
  ```bash
250
+ # Structure + manifest validation
251
  ..\\.venv\\Scripts\\openenv validate
252
 
253
+ # Determinism/reproducibility smoke test
254
+ ..\\.venv\\Scripts\\python scripts/reproducibility_check.py
255
+
256
+ # Submission-oriented local checks (without live inference)
257
+ bash scripts/pre_submit_validate.sh --skip-inference
258
+
259
+ # Build local image
260
+ docker build -t cloud-devops-env:phase1 -f Dockerfile .
261
+ ```
262
 
263
+ Optional local server:
 
264
 
265
+ ```bash
266
+ uvicorn server.app:app --host 0.0.0.0 --port 8000
 
267
  ```
268
 
269
+ ## Hugging Face Space Deployment
270
 
271
+ 1. Keep this front matter block intact (includes mandatory openenv tag).
272
+ 2. Push to Space (Docker SDK).
273
+ 3. Configure secrets/variables:
274
+ - HF_TOKEN
275
+ - API_BASE_URL (for example https://router.huggingface.co/v1)
276
+ - MODEL_NAME
277
+ 4. Wait for build completion.
278
+ 5. Verify:
279
+ - GET /health returns 200
280
+ - POST /reset returns 200
281
 
282
  Reference:
 
283
  - https://huggingface.co/docs/hub/spaces-config-reference
WEB_README.md CHANGED
@@ -1,172 +1,99 @@
1
- # Cloud DevOps RLEnv Web Playbook
2
 
3
- Use this page as your operator guide when stepping through incidents in /web.
4
 
5
- ## Mission
6
-
7
- You are the on-call SRE for a production outage.
8
-
9
- Your loop is always:
10
 
11
  1. Triage quickly.
12
- 2. Confirm root cause.
13
- 3. Apply the minimum safe fix.
14
- 4. Verify health and close.
15
-
16
- ## What Makes This Environment Non-Trivial
17
-
18
- - Large decoy inventory: many resources are intentionally irrelevant.
19
- - Ambiguous telemetry: medium/hard logs expose failing IPs, not only direct IDs.
20
- - Action cost: every step incurs a small penalty, so efficient plans outperform exploratory spam.
21
- - Hard-mode drift: delayed remediation causes additional system degradation.
22
-
23
- ## Command Reference
24
-
25
- ### list_resources
26
-
27
- - Purpose: enumerate all available entities.
28
- - Use first in most runs.
29
- - Required fields: command only.
30
-
31
- ### describe_resource
32
-
33
- - Purpose: inspect one resource configuration or current state.
34
- - Required fields: resource_id.
35
-
36
- ### view_logs
37
-
38
- - Purpose: read operational telemetry for one resource.
39
- - Required fields: resource_id.
40
-
41
- ### query_metadata
42
-
43
- - Purpose: resolve metadata such as ip_address -> resource_id.
44
- - Required fields: parameters.ip_address.
45
- - Typical medium/hard bridge action before mutation.
46
-
47
- ### update_security_group
48
 
49
- - Purpose: mutate network policy.
50
- - Required fields: resource_id, parameters.port, parameters.action.
51
- - Valid action values: allow, deny.
52
 
53
- ### restart_service
54
 
55
- - Purpose: restart a target service/instance.
56
- - Required fields: resource_id.
57
- - Use only after root-cause confirmation.
58
 
59
- ### submit_solution
 
60
 
61
- - Purpose: attempt to close incident.
62
- - Hard mode may continue with penalty if unresolved.
63
 
64
- ## JSON Action Examples
 
65
 
66
- ```json
67
- {"command":"list_resources"}
68
- ```
69
 
70
- ```json
71
- {"command":"view_logs","resource_id":"i-api"}
72
- ```
73
 
74
- ```json
75
- {"command":"query_metadata","parameters":{"ip_address":"10.0.4.5"}}
76
- ```
77
-
78
- ```json
79
- {"command":"update_security_group","resource_id":"sg-db","parameters":{"port":5432,"action":"allow"}}
80
- ```
81
-
82
- ```json
83
- {"command":"restart_service","resource_id":"i-web2"}
84
- ```
85
 
86
  ## Task Playbooks
87
 
88
- ### Easy: Web Access Recovery
89
 
90
  Objective:
91
-
92
- - Open port 80 on sg-web.
93
 
94
  Strong path:
 
 
95
 
96
- 1. list_resources
97
- 2. update_security_group(sg-web, port=80, action=allow)
98
-
99
- Optional safer path (slightly lower efficiency):
100
 
101
- 1. list_resources
102
- 2. describe_resource(sg-web)
103
- 3. update_security_group(sg-web, port=80, action=allow)
104
-
105
- ### Medium: API to DB Connectivity
106
 
107
  Objective:
108
-
109
- - Restore DB access by opening 5432 on sg-db, but only after diagnosis.
110
 
111
  Strong path:
112
-
113
- 1. list_resources
114
- 2. view_logs(i-api)
115
- 3. query_metadata(ip_address=10.0.4.5)
116
- 4. update_security_group(sg-db, port=5432, action=allow)
117
 
118
  Guardrail:
 
119
 
120
- - Mutating sg-db before logs + metadata lookup is penalized and does not resolve.
121
-
122
- ### Hard: Upstream Failure Under Drift
123
 
124
  Objective:
125
-
126
- - Find failing upstream from lb-main logs, resolve target identity, inspect, then restart i-web2.
127
 
128
  Strong path:
129
-
130
- 1. list_resources
131
- 2. view_logs(lb-main)
132
- 3. query_metadata(ip_address=10.0.8.22)
133
- 4. describe_resource(i-web2) or view_logs(i-web2)
134
- 5. restart_service(i-web2)
135
 
136
  Guardrails:
 
 
 
137
 
138
- - Restarting i-web1 is penalized.
139
- - Restarting i-web2 before investigation is penalized and blocked from resolution.
140
- - If unresolved after step 8, lb-external also fails (cascading failure).
141
-
142
- ## Rewards, Termination, And Score Semantics
143
 
144
- - Every action includes a small cost.
145
- - Discovery and correct remediation provide shaped positive rewards.
146
- - Unsafe or premature actions apply penalties.
147
- - Episode terminates on resolution or max step limit.
148
- - Submission score is clamped to a strict open interval (0,1).
149
- - success reflects actual incident resolution state.
150
 
151
- ## Observation Fields To Monitor
152
-
153
- - output: command result text
154
- - error: command-level error string
155
- - system_health_status: CRITICAL, DEGRADED, HEALTHY
156
- - reward: step reward
157
- - done: episode termination flag
158
- - metadata: task, step_count, resolved, achievements, action_cost
159
 
160
  ## Submission Contract Reminder
161
 
162
- inference.py must:
163
 
164
  - use OpenAI client
165
- - read API_BASE_URL, MODEL_NAME, HF_TOKEN
166
- - emit strict stdout lines with START, STEP, END markers
167
-
168
- Practical strategy:
169
-
170
- - Inspect first, mutate second.
171
- - Use query_metadata whenever logs give only IP clues.
172
- - Prefer minimal actions for higher score.
 
1
+ # Cloud DevOps RLEnv
2
 
3
+ Use this page as the operator playbook in the /web interface.
4
 
5
+ ## Core Loop
 
 
 
 
6
 
7
  1. Triage quickly.
8
+ 2. Build evidence from logs and metadata.
9
+ 3. Apply the minimum safe remediation.
10
+ 4. Verify resolution and stop.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
11
 
12
+ Why this matters: each action has a cost (`-0.01`), so shorter correct trajectories score higher.
 
 
13
 
14
+ ## Command Cheat Sheet
15
 
16
+ - `list_resources`
17
+ : Start here to map real targets vs decoys.
 
18
 
19
+ - `describe_resource(resource_id)`
20
+ : Inspect config/state details.
21
 
22
+ - `view_logs(resource_id)`
23
+ : Primary source of root-cause evidence.
24
 
25
+ - `query_metadata(parameters={"ip_address": "..."})`
26
+ : Resolve IP-only clues to resource IDs (mandatory multi-hop step in medium/hard).
27
 
28
+ - `update_security_group(resource_id, parameters)`
29
+ : Requires `parameters.port` and `parameters.action` where action is `allow` or `deny`.
 
30
 
31
+ - `restart_service(resource_id)`
32
+ : Use only after causal evidence confirms target.
 
33
 
34
+ - `submit_solution`
35
+ : Useful for explicit closure checks; unresolved hard submissions are penalized and continue.
 
 
 
 
 
 
 
 
 
36
 
37
  ## Task Playbooks
38
 
39
+ ### Easy
40
 
41
  Objective:
42
+ - Restore web access by allowing port `80` on `sg-web`.
 
43
 
44
  Strong path:
45
+ 1. `list_resources`
46
+ 2. `update_security_group("sg-web", {"port": 80, "action": "allow"})`
47
 
48
+ Typical outcome:
49
+ - 2 steps, score around `0.78`
 
 
50
 
51
+ ### Medium
 
 
 
 
52
 
53
  Objective:
54
+ - Restore API to DB connectivity by allowing port `5432` on `sg-db`.
 
55
 
56
  Strong path:
57
+ 1. `list_resources`
58
+ 2. `view_logs("i-api")`
59
+ 3. `query_metadata({"ip_address": "10.0.4.5"})`
60
+ 4. `update_security_group("sg-db", {"port": 5432, "action": "allow"})`
 
61
 
62
  Guardrail:
63
+ - Applying SG changes before logs + metadata lookup is penalized and does not resolve the incident.
64
 
65
+ ### Hard
 
 
66
 
67
  Objective:
68
+ - Recover checkout flow by identifying the failing upstream and restarting `i-web2` safely.
 
69
 
70
  Strong path:
71
+ 1. `list_resources`
72
+ 2. `view_logs("lb-main")`
73
+ 3. `query_metadata({"ip_address": "10.0.8.22"})`
74
+ 4. `describe_resource("i-web2")` or `view_logs("i-web2")`
75
+ 5. `restart_service("i-web2")`
 
76
 
77
  Guardrails:
78
+ - Restarting `i-web1` is penalized.
79
+ - Restarting `i-web2` without investigation is penalized.
80
+ - If unresolved after step 8, `lb-external` also fails (cascading failure).
81
 
82
+ ## Reading Metadata
 
 
 
 
83
 
84
+ Watch these response fields each step:
 
 
 
 
 
85
 
86
+ - `system_health_status`: `CRITICAL` / `DEGRADED` / `HEALTHY`
87
+ - `done`: episode ended or still running
88
+ - `reward`: immediate signal after action cost and shaping
89
+ - `metadata.resolved`: authoritative success flag
90
+ - `metadata.termination_reason`: why episode ended
91
+ - `metadata.reward_breakdown`: transparent reward events for grader/debug inspection
 
 
92
 
93
  ## Submission Contract Reminder
94
 
95
+ `inference.py` must:
96
 
97
  - use OpenAI client
98
+ - read `API_BASE_URL`, `MODEL_NAME`, `HF_TOKEN`
99
+ - emit strict stdout markers: `[START]`, `[STEP]`, `[END]`
 
 
 
 
 
 
scripts/reproducibility_check.py ADDED
@@ -0,0 +1,125 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python
2
+ """Deterministic reproducibility smoke test for Cloud DevOps RLEnv."""
3
+
4
+ from __future__ import annotations
5
+
6
+ import asyncio
7
+ import json
8
+ import sys
9
+ from pathlib import Path
10
+ from typing import Any
11
+
12
+ REPO_ROOT = Path(__file__).resolve().parent.parent
13
+ if str(REPO_ROOT) not in sys.path:
14
+ sys.path.insert(0, str(REPO_ROOT))
15
+
16
+ from env import CloudDevOpsEnv
17
+ from models import CloudAction
18
+
19
+ SCORE_MIN = 0.001
20
+ SCORE_MAX = 0.999
21
+
22
+ # Fixed trajectories that should always resolve the incidents.
23
+ POLICY_BY_TASK: dict[str, list[dict[str, Any]]] = {
24
+ "easy": [
25
+ {"command": "list_resources"},
26
+ {
27
+ "command": "update_security_group",
28
+ "resource_id": "sg-web",
29
+ "parameters": {"port": 80, "action": "allow"},
30
+ },
31
+ ],
32
+ "medium": [
33
+ {"command": "list_resources"},
34
+ {"command": "view_logs", "resource_id": "i-api"},
35
+ {
36
+ "command": "query_metadata",
37
+ "parameters": {"ip_address": "10.0.4.5"},
38
+ },
39
+ {
40
+ "command": "update_security_group",
41
+ "resource_id": "sg-db",
42
+ "parameters": {"port": 5432, "action": "allow"},
43
+ },
44
+ ],
45
+ "hard": [
46
+ {"command": "list_resources"},
47
+ {"command": "view_logs", "resource_id": "lb-main"},
48
+ {
49
+ "command": "query_metadata",
50
+ "parameters": {"ip_address": "10.0.8.22"},
51
+ },
52
+ {"command": "describe_resource", "resource_id": "i-web2"},
53
+ {"command": "restart_service", "resource_id": "i-web2"},
54
+ ],
55
+ }
56
+
57
+
58
+ async def run_policy(task_name: str) -> dict[str, Any]:
59
+ env = CloudDevOpsEnv(task_name=task_name)
60
+ await env.reset()
61
+
62
+ trajectory: list[dict[str, Any]] = []
63
+ rewards: list[float] = []
64
+ last = None
65
+
66
+ try:
67
+ for index, raw_action in enumerate(POLICY_BY_TASK[task_name], start=1):
68
+ action = CloudAction(**raw_action)
69
+ result = await env.step(action)
70
+ rewards.append(float(result.reward))
71
+ trajectory.append(
72
+ {
73
+ "step": index,
74
+ "command": action.command,
75
+ "resource_id": action.resource_id,
76
+ "reward": round(float(result.reward), 4),
77
+ "done": bool(result.done),
78
+ "error": result.observation.error,
79
+ "status": result.observation.system_health_status,
80
+ "resolved": bool(result.info.get("resolved", False)),
81
+ }
82
+ )
83
+ last = result
84
+ if result.done:
85
+ break
86
+
87
+ if last is None:
88
+ raise RuntimeError(f"Task {task_name} produced no steps")
89
+
90
+ score = max(SCORE_MIN, min(sum(rewards), SCORE_MAX))
91
+ return {
92
+ "task": task_name,
93
+ "resolved": bool(last.info.get("resolved", False)),
94
+ "steps": len(trajectory),
95
+ "score": round(score, 3),
96
+ "trajectory": trajectory,
97
+ }
98
+ finally:
99
+ await env.close()
100
+
101
+
102
+ async def main() -> None:
103
+ summary: dict[str, Any] = {}
104
+
105
+ for task_name in ("easy", "medium", "hard"):
106
+ run_1 = await run_policy(task_name)
107
+ run_2 = await run_policy(task_name)
108
+
109
+ if run_1["trajectory"] != run_2["trajectory"]:
110
+ raise SystemExit(f"Determinism check failed for task={task_name}: trajectories differ")
111
+ if not run_1["resolved"]:
112
+ raise SystemExit(f"Policy failed to resolve task={task_name}")
113
+
114
+ summary[task_name] = {
115
+ "steps": run_1["steps"],
116
+ "score": run_1["score"],
117
+ "resolved": run_1["resolved"],
118
+ }
119
+
120
+ print("Deterministic reproducibility check passed")
121
+ print(json.dumps(summary, indent=2, sort_keys=True))
122
+
123
+
124
+ if __name__ == "__main__":
125
+ asyncio.run(main())
server/cloud_devops_env_environment.py CHANGED
@@ -217,6 +217,20 @@ class CloudDevopsEnvironment(Environment):
217
  self._achievements.add(achievement)
218
  return points
219
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
220
  def reset(self) -> CloudObservation: # type: ignore[override]
221
  """Reset the environment to the initial state for the selected task."""
222
  self._achievements.clear()
@@ -242,6 +256,11 @@ class CloudDevopsEnvironment(Environment):
242
  "resolved": False,
243
  "task": self.task_name,
244
  "total_resources": len(self._state_data.resources),
 
 
 
 
 
245
  },
246
  echoed_message="Cloud Devops Env environment ready!",
247
  message_length=0,
@@ -257,9 +276,20 @@ class CloudDevopsEnvironment(Environment):
257
 
258
  state.step_count += 1
259
  reward = -self.ACTION_COST
 
 
 
260
  done = False
261
  output = ""
262
  error = None
 
 
 
 
 
 
 
 
263
 
264
  try:
265
  if action.command == "list_resources":
@@ -276,11 +306,14 @@ class CloudDevopsEnvironment(Environment):
276
  output = str(state.resources[action.resource_id])
277
 
278
  if self.task_name == "easy" and action.resource_id == "sg-web":
279
- reward += self._reward_once("read_sg", 0.2)
280
  elif self.task_name == "medium" and action.resource_id == "sg-db":
281
- reward += self._reward_once("read_sg", 0.2)
282
  elif self.task_name == "hard" and action.resource_id == "i-web2":
283
- reward += self._reward_once("inspect_target", 0.2)
 
 
 
284
 
285
  elif action.command == "view_logs":
286
  if not action.resource_id:
@@ -293,11 +326,14 @@ class CloudDevopsEnvironment(Environment):
293
  output = str(res.get("logs", "No logs available for this resource."))
294
 
295
  if self.task_name == "medium" and action.resource_id == "i-api":
296
- reward += self._reward_once("read_logs", 0.2)
297
  elif self.task_name == "hard" and action.resource_id == "lb-main":
298
- reward += self._reward_once("inspect_lb", 0.2)
299
  elif self.task_name == "hard" and action.resource_id == "i-web2":
300
- reward += self._reward_once("inspect_target", 0.2)
 
 
 
301
 
302
  elif action.command == "query_metadata":
303
  ip_address = None
@@ -314,9 +350,15 @@ class CloudDevopsEnvironment(Environment):
314
 
315
  output = f"Metadata lookup: ip_address={ip_address} resource_id={resource_id}"
316
  if self.task_name == "medium" and str(ip_address) == "10.0.4.5":
317
- reward += self._reward_once("lookup_db_target", 0.2)
 
 
 
318
  elif self.task_name == "hard" and str(ip_address) == "10.0.8.22":
319
- reward += self._reward_once("lookup_upstream_target", 0.2)
 
 
 
320
 
321
  elif action.command == "update_security_group":
322
  if not action.resource_id:
@@ -350,8 +392,9 @@ class CloudDevopsEnvironment(Environment):
350
  and rule_action == "allow"
351
  ):
352
  state.is_resolved = True
353
- reward += 0.8
354
  done = True
 
355
  output += "\nSUCCESS: Web server is now accessible!"
356
  elif (
357
  self.task_name == "medium"
@@ -365,17 +408,18 @@ class CloudDevopsEnvironment(Environment):
365
  )
366
  if investigated:
367
  state.is_resolved = True
368
- reward += 0.6
369
  done = True
 
370
  output += "\nSUCCESS: Database connection restored!"
371
  else:
372
- reward -= 0.1
373
  output += (
374
  "\nWARNING: Change applied without incident triage. "
375
  "Inspect API logs and resolve DB IP via query_metadata before closing the incident."
376
  )
377
  elif rule_action == "deny":
378
- reward -= 0.1
379
  output += "\nWARNING: Deny rule applied during outage remediation."
380
 
381
  elif action.command == "restart_service":
@@ -399,17 +443,18 @@ class CloudDevopsEnvironment(Environment):
399
  "logs"
400
  ] = "INFO: Restart successful. Memory cleared."
401
  state.is_resolved = True
402
- reward += 0.8
403
  done = True
 
404
  output += "\nSUCCESS: OutOfMemory loop broken. System stable."
405
  else:
406
- reward -= 0.1
407
  output += (
408
  "\nWARNING: Restart denied by change policy. "
409
  "Find failing upstream IP from lb-main, resolve it with query_metadata, and inspect i-web2 first."
410
  )
411
  elif action.resource_id == "i-web1":
412
- reward -= 0.2
413
  output += (
414
  "\nWARNING: You restarted a healthy production server! "
415
  "Users dropped."
@@ -418,18 +463,20 @@ class CloudDevopsEnvironment(Environment):
418
  elif action.command == "submit_solution":
419
  if state.is_resolved:
420
  done = True
 
421
  output = "Solution verified. System is HEALTHY."
422
  else:
423
  if self.task_name == "hard":
424
  # In hard mode, unresolved submission should not abort the run.
425
  done = False
426
- reward -= 0.1
427
  output = (
428
  "Solution incorrect. Incident is still CRITICAL. "
429
  "Continue triage and remediation before submitting."
430
  )
431
  else:
432
  done = True
 
433
  output = "Solution incorrect. System is still CRITICAL."
434
 
435
  else:
@@ -440,16 +487,26 @@ class CloudDevopsEnvironment(Environment):
440
  output = f"Command Failed: {error}"
441
 
442
  cascade_penalty, cascade_msg = self._apply_cascading_failure()
443
- reward += cascade_penalty
444
  if cascade_msg:
445
  output = f"{output}{cascade_msg}" if output else cascade_msg.strip()
446
 
447
  if state.step_count >= self.MAX_STEPS and not done:
448
  done = True
 
449
  timeout_suffix = "\nTIMEOUT: Max steps reached."
450
  output = f"{output}{timeout_suffix}" if output else timeout_suffix.strip()
451
 
 
452
  reward = max(-1.0, min(1.0, reward))
 
 
 
 
 
 
 
 
453
  lb_external = state.resources.get("lb-external", {})
454
  if state.is_resolved:
455
  status = "HEALTHY"
@@ -464,6 +521,11 @@ class CloudDevopsEnvironment(Environment):
464
  "achievements": sorted(self._achievements),
465
  "total_resources": len(state.resources),
466
  "action_cost": self.ACTION_COST,
 
 
 
 
 
467
  }
468
 
469
  return CloudObservation(
 
217
  self._achievements.add(achievement)
218
  return points
219
 
220
+ def _task_objective(self) -> str:
221
+ objectives = {
222
+ "easy": "Restore web access by allowing port 80 on sg-web.",
223
+ "medium": (
224
+ "Restore API to DB connectivity by reading i-api logs, resolving DB IP via "
225
+ "query_metadata, then allowing port 5432 on sg-db."
226
+ ),
227
+ "hard": (
228
+ "Recover checkout path by tracing lb-main upstream IP, resolving it with "
229
+ "query_metadata, inspecting i-web2, and restarting i-web2 safely."
230
+ ),
231
+ }
232
+ return objectives[self.task_name]
233
+
234
  def reset(self) -> CloudObservation: # type: ignore[override]
235
  """Reset the environment to the initial state for the selected task."""
236
  self._achievements.clear()
 
256
  "resolved": False,
257
  "task": self.task_name,
258
  "total_resources": len(self._state_data.resources),
259
+ "objective": self._task_objective(),
260
+ "deterministic": True,
261
+ "max_steps": self.MAX_STEPS,
262
+ "action_cost": self.ACTION_COST,
263
+ "hard_cascade_trigger_step": 8,
264
  },
265
  echoed_message="Cloud Devops Env environment ready!",
266
  message_length=0,
 
276
 
277
  state.step_count += 1
278
  reward = -self.ACTION_COST
279
+ reward_breakdown: list[dict[str, object]] = [
280
+ {"event": "action_cost", "delta": -self.ACTION_COST}
281
+ ]
282
  done = False
283
  output = ""
284
  error = None
285
+ termination_reason = "in_progress"
286
+
287
+ def add_reward(delta: float, event: str) -> None:
288
+ nonlocal reward
289
+ if abs(delta) < 1e-12:
290
+ return
291
+ reward += delta
292
+ reward_breakdown.append({"event": event, "delta": round(float(delta), 4)})
293
 
294
  try:
295
  if action.command == "list_resources":
 
306
  output = str(state.resources[action.resource_id])
307
 
308
  if self.task_name == "easy" and action.resource_id == "sg-web":
309
+ add_reward(self._reward_once("read_sg", 0.2), "inspect_web_sg")
310
  elif self.task_name == "medium" and action.resource_id == "sg-db":
311
+ add_reward(self._reward_once("read_sg", 0.2), "inspect_db_sg")
312
  elif self.task_name == "hard" and action.resource_id == "i-web2":
313
+ add_reward(
314
+ self._reward_once("inspect_target", 0.2),
315
+ "inspect_target_instance",
316
+ )
317
 
318
  elif action.command == "view_logs":
319
  if not action.resource_id:
 
326
  output = str(res.get("logs", "No logs available for this resource."))
327
 
328
  if self.task_name == "medium" and action.resource_id == "i-api":
329
+ add_reward(self._reward_once("read_logs", 0.2), "inspect_api_logs")
330
  elif self.task_name == "hard" and action.resource_id == "lb-main":
331
+ add_reward(self._reward_once("inspect_lb", 0.2), "inspect_lb_logs")
332
  elif self.task_name == "hard" and action.resource_id == "i-web2":
333
+ add_reward(
334
+ self._reward_once("inspect_target", 0.2),
335
+ "inspect_target_logs",
336
+ )
337
 
338
  elif action.command == "query_metadata":
339
  ip_address = None
 
350
 
351
  output = f"Metadata lookup: ip_address={ip_address} resource_id={resource_id}"
352
  if self.task_name == "medium" and str(ip_address) == "10.0.4.5":
353
+ add_reward(
354
+ self._reward_once("lookup_db_target", 0.2),
355
+ "resolve_db_ip_dependency",
356
+ )
357
  elif self.task_name == "hard" and str(ip_address) == "10.0.8.22":
358
+ add_reward(
359
+ self._reward_once("lookup_upstream_target", 0.2),
360
+ "resolve_upstream_ip_dependency",
361
+ )
362
 
363
  elif action.command == "update_security_group":
364
  if not action.resource_id:
 
392
  and rule_action == "allow"
393
  ):
394
  state.is_resolved = True
395
+ add_reward(0.8, "resolve_easy_web_ingress")
396
  done = True
397
+ termination_reason = "resolved_easy"
398
  output += "\nSUCCESS: Web server is now accessible!"
399
  elif (
400
  self.task_name == "medium"
 
408
  )
409
  if investigated:
410
  state.is_resolved = True
411
+ add_reward(0.6, "resolve_medium_db_connectivity")
412
  done = True
413
+ termination_reason = "resolved_medium"
414
  output += "\nSUCCESS: Database connection restored!"
415
  else:
416
+ add_reward(-0.1, "unsafe_change_without_triage")
417
  output += (
418
  "\nWARNING: Change applied without incident triage. "
419
  "Inspect API logs and resolve DB IP via query_metadata before closing the incident."
420
  )
421
  elif rule_action == "deny":
422
+ add_reward(-0.1, "deny_rule_during_incident")
423
  output += "\nWARNING: Deny rule applied during outage remediation."
424
 
425
  elif action.command == "restart_service":
 
443
  "logs"
444
  ] = "INFO: Restart successful. Memory cleared."
445
  state.is_resolved = True
446
+ add_reward(0.8, "resolve_hard_upstream_recovery")
447
  done = True
448
+ termination_reason = "resolved_hard"
449
  output += "\nSUCCESS: OutOfMemory loop broken. System stable."
450
  else:
451
+ add_reward(-0.1, "restart_without_root_cause")
452
  output += (
453
  "\nWARNING: Restart denied by change policy. "
454
  "Find failing upstream IP from lb-main, resolve it with query_metadata, and inspect i-web2 first."
455
  )
456
  elif action.resource_id == "i-web1":
457
+ add_reward(-0.2, "restart_healthy_node")
458
  output += (
459
  "\nWARNING: You restarted a healthy production server! "
460
  "Users dropped."
 
463
  elif action.command == "submit_solution":
464
  if state.is_resolved:
465
  done = True
466
+ termination_reason = "resolved_submit_solution"
467
  output = "Solution verified. System is HEALTHY."
468
  else:
469
  if self.task_name == "hard":
470
  # In hard mode, unresolved submission should not abort the run.
471
  done = False
472
+ add_reward(-0.1, "premature_submit_hard")
473
  output = (
474
  "Solution incorrect. Incident is still CRITICAL. "
475
  "Continue triage and remediation before submitting."
476
  )
477
  else:
478
  done = True
479
+ termination_reason = "incorrect_submit"
480
  output = "Solution incorrect. System is still CRITICAL."
481
 
482
  else:
 
487
  output = f"Command Failed: {error}"
488
 
489
  cascade_penalty, cascade_msg = self._apply_cascading_failure()
490
+ add_reward(cascade_penalty, "cascading_failure_penalty")
491
  if cascade_msg:
492
  output = f"{output}{cascade_msg}" if output else cascade_msg.strip()
493
 
494
  if state.step_count >= self.MAX_STEPS and not done:
495
  done = True
496
+ termination_reason = "max_steps_timeout"
497
  timeout_suffix = "\nTIMEOUT: Max steps reached."
498
  output = f"{output}{timeout_suffix}" if output else timeout_suffix.strip()
499
 
500
+ raw_reward = reward
501
  reward = max(-1.0, min(1.0, reward))
502
+ if reward != raw_reward:
503
+ reward_breakdown.append(
504
+ {
505
+ "event": "reward_clip",
506
+ "delta": round(float(reward - raw_reward), 4),
507
+ }
508
+ )
509
+
510
  lb_external = state.resources.get("lb-external", {})
511
  if state.is_resolved:
512
  status = "HEALTHY"
 
521
  "achievements": sorted(self._achievements),
522
  "total_resources": len(state.resources),
523
  "action_cost": self.ACTION_COST,
524
+ "objective": self._task_objective(),
525
+ "deterministic": True,
526
+ "max_steps": self.MAX_STEPS,
527
+ "termination_reason": termination_reason if done else "in_progress",
528
+ "reward_breakdown": reward_breakdown,
529
  }
530
 
531
  return CloudObservation(