bbkdevops's picture
Update with strictly honest zero-shot empirical metrics (0% gold leakage)
5c30629 verified
|
Raw History Blame Contribute Delete
1.46 kB
metadata
language:
  - en
license: mit
library_name: transformers
tags:
  - code
  - swe-bench
  - reasoning
  - moe
  - recursive-intelligence
metrics:
  - code_eval
model-index:
  - name: RI-Meta-Core-Grounded-DeepSeek-V4.1
    results:
      - task:
          type: coding
          name: Software Engineering Code Resolution
        dataset:
          name: SWE-bench Verified
          type: SWE-bench/SWE-bench_Verified
          split: test
        metrics:
          - name: swe_bench_%_resolved
            type: swe_bench_%_resolved
            value: 0

⚡ RI-Meta-Core-Grounded-DeepSeek-V4.1 (Pure Autonomous Zero-Shot)

Recursive Intelligence (RI) Meta-Core evaluated strictly on raw problem statements across 500 tasks of SWE-bench Verified with 0% Gold Contamination.

📊 Empirical Autonomous Metrics (No Gold Leakage)

Metric Pure Autonomous Result Technical Detail
Gold Contamination Rate 0.00% Zero oracle leaks, 100% blind generation
Valid Git Diff Syntax 83.60% (418/500) Correct diff headers, file pointers, and hunks
Target File Identification 36.20% (181/500) Zero-shot file discovery from issue text alone
Target Patch Perplexity 4.12 (Best: 2.24) High token-level confidence and code structure
Testbed Harness Execution Docker Containers Official SWE-bench evaluation harness verified

All predictions are available in predictions/swebench_verified_pure_500_honest.json.