RapidRAW Nonlocal RAW Denoise — runtime artifacts

This repository distributes runtime conversion artifacts for the RapidRAW Nonlocal Bayer RAW denoiser. It is not a newly trained model: both artifacts are converted, without retraining or precision changes, from the same pinned upstream checkpoint.

RapidRAW installs and verifies these files automatically (MCP tool install_model with kind: "nonlocal"; see the denoise README). Earlier revisions of this repository remain available at their commits, and RapidRAW pins one immutable revision.

Upstream attribution and license

  • Project: MIA-UIB/nonlocal-matchfilter
  • Paper: “Learned Nonlocal Feature Matching and Filtering for RAW Image Denoising”, Marco Sánchez-Beeckman and Antoni Buades, arXiv:2604.17453.
  • Source checkpoint: upstream v1.0.0 release weights (rawnoise_25_15_9nbr.ckpt), tensor-only SHA-256 c16747d852b93a95908792cdbac901f89cca98e35b91ea7db6214de42fbd3cad.
  • License: MIT (see LICENSE; copyright Marco Sánchez Beeckman). Attribution retained per the license terms.

Input contract (both variants)

  • RAW Bayer input packed to 4 RGBG planes plus a 4-channel noise-level map: raw_with_noise [1, 8, 320, 320], float32.
  • Output: denoised packed RAW denoised_raw [1, 4, 320, 320], float32.
  • RapidRAW's host tiling (not a property of the model): 320 px packed tiles with a 40 px halo (240 px retained core), one pass per tile. A 32.5 MP photograph takes 150 tiles.

Variants

Variant Backend File SHA-256
CoreML (macOS, direct CoreML.framework) native-coreml-v1 coreml/packed.mlpackage ecf8f41b72b8e79a7210f61b13b3a21ad2ddbc79c475039be367fc011a98f5d3
ONNX (CPU / Linux CUDA) native-onnx-v1 onnx/model.onnx ba788a97f247058ea03d14bbf2f289adc8f4666039f27d0d530e6f67d23dd73c
  • CoreML is a direct TorchScript conversion (coremltools-direct-torchscript, FLOAT32, static-packed sampler) for native CoreML.framework inference on macOS. It is not an ONNX/CoreML-execution-provider artifact. It is unchanged from the previous revision.
  • ONNX is the FLOAT32 export of the same checkpoint with its graph rewritten for speed: the all-zero bias nodes are removed and the dense and grouped 1×1 convolutions are expressed as matrix multiplications. Weights and precision are unchanged; results can differ from the unrewritten graph at rounding level because the matrix multiplications sum in a different order. onnx/rewrite-report.json records the rewrite and onnx/manifest.json records the SHA-256 of the unrewritten export it was derived from (df7f21ddbdfecd0984896a902623f75aa883d25b5f862c49d8c2b3f63a7dd2d9).
  • distribution.json lists every downloadable file with size and SHA-256. RapidRAW downloads by immutable commit revision and verifies distribution.json, every file, and the inner bundle contracts — never main alone.

Measured evidence and limits

  • ONNX graph rewrite: the rewritten graph passes the same frozen fixture gates as the unrewritten export on all 18 fixture tiles, on both CPU and CUDA (elementwise |difference| ≤ 1e-4 + 1e-4·|reference|, mean absolute error ≤ 1e-5, 99th-percentile absolute error ≤ 1e-4). On an RTX 5060 Ti with RapidRAW's runtime settings (CUDA, TF32 and graph optimization off), the time per 320×320 tile falls from 370 ms to 153 ms (2.4×). CPU time was not measured, and the CoreML package is unchanged.
  • Tiling: against a converged large-halo reference on nine 768-pixel crops (three photographs, three crops each), the 40 px halo differs by at most 0.031 noise standard deviations (99.9% within 0.012). In 8-bit sRGB test renders of the same three photographs (AHD demosaic, camera white balance, no tone edits), 99.9% of pixels differ from the previous 64 px halo by at most 0.71 of 255 levels at the model's normal noise level (worst single pixel 6.8 levels). With the noise-level multiplier halved, a research setting RapidRAW does not expose, differences are larger: 99.9% within 2.2 levels, about 0.1% of pixels above 2 levels, worst single pixel 44 levels. RapidRAW's Intensity setting blends the result with the original and scales these differences down with it. These are numeric comparisons.
  • Speed: at the measured 153 ms per tile, a 32.5 MP photograph needs about 23 s of model time with the 40 px halo (150 tiles) against about 38 s with the 64 px halo (247 tiles). RAW decoding, noise estimation and DNG writing are not included.
  • CoreML: 18/18 fixture tiles pass through RapidRAW's CoreML backend. Tile results do not depend on the halo; full-photograph CoreML parity was last measured with the previous 64 px halo and has not been re-run.
  • Experimental: commercial parity has not been established.

Intended use

  • Automatic installation by RapidRAW.
  • No hosted inference API is provided or expected.
Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for sheldonxxxx/RapidRAW-Nonlocal-Denoise