RapidRAW Nonlocal RAW Denoise — runtime artifacts
This repository distributes runtime conversion artifacts for the
RapidRAW Nonlocal Bayer RAW
denoiser. It is not a newly trained model: both artifacts are converted,
without retraining or precision changes, from the same pinned upstream
checkpoint.
RapidRAW installs and verifies these files automatically (MCP tool
install_model with kind: "nonlocal"; see the
denoise README).
Earlier revisions of this repository remain available at their commits, and
RapidRAW pins one immutable revision.
Upstream attribution and license
- Project: MIA-UIB/nonlocal-matchfilter
- Paper: “Learned Nonlocal Feature Matching and Filtering for RAW Image
Denoising”, Marco Sánchez-Beeckman and Antoni Buades, arXiv:2604.17453.
- Source checkpoint: upstream
v1.0.0 release weights
(rawnoise_25_15_9nbr.ckpt), tensor-only SHA-256
c16747d852b93a95908792cdbac901f89cca98e35b91ea7db6214de42fbd3cad.
- License: MIT (see
LICENSE; copyright Marco Sánchez Beeckman).
Attribution retained per the license terms.
Input contract (both variants)
- RAW Bayer input packed to 4 RGBG planes plus a 4-channel noise-level map:
raw_with_noise [1, 8, 320, 320], float32.
- Output: denoised packed RAW
denoised_raw [1, 4, 320, 320], float32.
- RapidRAW's host tiling (not a property of the model): 320 px packed tiles
with a 40 px halo (240 px retained core), one pass per tile. A 32.5 MP
photograph takes 150 tiles.
Variants
| Variant |
Backend |
File |
SHA-256 |
| CoreML (macOS, direct CoreML.framework) |
native-coreml-v1 |
coreml/packed.mlpackage |
ecf8f41b72b8e79a7210f61b13b3a21ad2ddbc79c475039be367fc011a98f5d3 |
| ONNX (CPU / Linux CUDA) |
native-onnx-v1 |
onnx/model.onnx |
ba788a97f247058ea03d14bbf2f289adc8f4666039f27d0d530e6f67d23dd73c |
- CoreML is a direct TorchScript conversion (
coremltools-direct-torchscript,
FLOAT32, static-packed sampler) for native CoreML.framework inference on
macOS. It is not an ONNX/CoreML-execution-provider artifact. It is
unchanged from the previous revision.
- ONNX is the FLOAT32 export of the same checkpoint with its graph rewritten
for speed: the all-zero bias nodes are removed and the dense and grouped
1×1 convolutions are expressed as matrix multiplications. Weights and
precision are unchanged; results can differ from the unrewritten graph at
rounding level because the matrix multiplications sum in a different order.
onnx/rewrite-report.json records the rewrite and onnx/manifest.json
records the SHA-256 of the unrewritten export it was derived from
(df7f21ddbdfecd0984896a902623f75aa883d25b5f862c49d8c2b3f63a7dd2d9).
distribution.json lists every downloadable file with size and SHA-256.
RapidRAW downloads by immutable commit revision and verifies
distribution.json, every file, and the inner bundle contracts — never
main alone.
Measured evidence and limits
- ONNX graph rewrite: the rewritten graph passes the same frozen fixture
gates as the unrewritten export on all 18 fixture tiles, on both CPU and
CUDA (elementwise |difference| ≤ 1e-4 + 1e-4·|reference|, mean absolute
error ≤ 1e-5, 99th-percentile absolute error ≤ 1e-4). On an RTX 5060 Ti
with RapidRAW's runtime settings (CUDA, TF32 and graph optimization off),
the time per 320×320 tile falls from 370 ms to 153 ms (2.4×). CPU time was
not measured, and the CoreML package is unchanged.
- Tiling: against a converged large-halo reference on nine 768-pixel crops
(three photographs, three crops each), the 40 px halo differs by at most
0.031 noise standard deviations (99.9% within 0.012). In 8-bit sRGB test
renders of the same three photographs (AHD demosaic, camera white balance,
no tone edits), 99.9% of pixels differ from the previous 64 px halo by at
most 0.71 of 255 levels at the model's normal noise level (worst single
pixel 6.8 levels). With the noise-level multiplier halved, a research
setting RapidRAW does not expose, differences are larger: 99.9% within 2.2
levels, about 0.1% of pixels above 2 levels, worst single pixel 44 levels.
RapidRAW's Intensity setting blends the result with the original and scales
these differences down with it. These are numeric comparisons.
- Speed: at the measured 153 ms per tile, a 32.5 MP photograph needs about
23 s of model time with the 40 px halo (150 tiles) against about 38 s with
the 64 px halo (247 tiles). RAW decoding, noise estimation and DNG writing
are not included.
- CoreML: 18/18 fixture tiles pass through RapidRAW's CoreML backend. Tile
results do not depend on the halo; full-photograph CoreML parity was last
measured with the previous 64 px halo and has not been re-run.
- Experimental: commercial parity has not been established.
Intended use
- Automatic installation by RapidRAW.
- No hosted inference API is provided or expected.