# Browser image upscaler: model licence gate and pinned inventory

Route A: official **Real-ESRGAN realesr-general-x4v3**, SRVGGNetCompact (64 features, 32 body convolutions, PReLU, native x4), converted from the official release to ONNX ourselves. BSD-3-Clause covers the upstream project and its distributed weights; there is no separate restricted weights licence in this release. The official release is https://github.com/xinntao/Real-ESRGAN/releases/tag/v0.2.5.0 . The source repository revision inspected is `a4abfb2979a7bbff3f69f58f58ae324608821e27`.

| Candidate | Licence evidence | Decision |
| --- | --- | --- |
| Official realesr-general-x4v3 | Actual upstream LICENSE: BSD-3-Clause, retained verbatim below and in LICENSE-Real-ESRGAN.txt; official release weights | Selected; raw model output by default, optional Soft blend (quality evidence below) |
| @upscalerjs/esrgan-slim 1.0.0 / esrgan-medium 1.0.0 | Each actual package.json licence field is `MIT`, and each archive contains package/LICENSE headed “MIT License”, copyright (c) 2022 Kevin Scott, with the MIT permission grant | Evaluated; no redistribution, raw quality below bicubic on the same fixture |
| UpscalerJS / TensorFlow.js route B | Not selected or redistributed | No second ML runtime needed |

No third-party ONNX claim is relied on. `scripts/convert_upscaler.py` is the exact conversion script. Use its documented throwaway venv, torch 2.5.1+cpu and onnx 1.17.0; the venv and original checkpoint are never shipped. Source pins:

| Source | URL | SHA-256 | Bytes |
| --- | --- | --- | --- |
| Official checkpoint | https://github.com/xinntao/Real-ESRGAN/releases/download/v0.2.5.0/realesr-general-x4v3.pth | `8dc7edb9ac80ccdc30c3a5dca6616509367f05fbc184ad95b731f05bece96292` | 4,885,111 |
| Official architecture | https://raw.githubusercontent.com/xinntao/Real-ESRGAN/a4abfb2979a7bbff3f69f58f58ae324608821e27/realesrgan/archs/srvgg_arch.py | `7419a75e789ca09295266563d74a6d7c08d6c24502af6b0a919b1811d0a0be59` | 2,722 |

Export: ONNX opset **17**, input **input** `[1,3,height,width]` float32 RGB / 255; output **output** `[1,3,height*4,width*4]` float32 RGB. Dynamic spatial dimensions, batch one. Convolution, PReLU, DepthToSpace, nearest Resize and Add use the existing **onnxruntime-web 1.22.0 WASM** in a single-thread worker. No quantization; no before/after quantization PSNR is applicable. No model weight interpolation. Outputs are clamped to [0,255]. Sharp, the shipped default, returns the raw model RGB output; Soft optionally combines it with ordinary bicubic in a 50/50 RGB blend. This is output processing, not model training or weight interpolation. Real Chromium inference, PSNR, SSIM and mean absolute Laplacian of luma are measured at test time in tests/test_image_upscaler.py against a deterministic PIL-generated picture and PIL bicubic baseline.

All new binaries are under 25 MB and supplied as ordinary repository files, left uncommitted in this working tree. No new fetch_models.py entry or model gitignore is needed. assets.json is the worker's new model inventory; shared-assets.json records already-vendored dependencies without copying them. Those original JSON inventory files and this document are metadata, not third-party vendor binaries.

| New file | Source URL / reproducible conversion | SHA-256 | Bytes | Licence |
| --- | --- | --- | --- | --- |
| realesr-general-x4v3.onnx | https://github.com/xinntao/Real-ESRGAN/releases/download/v0.2.5.0/realesr-general-x4v3.pth#conversion=scripts/convert_upscaler.py;opset=17 | `83eced3ba3824e4211a261fd7862c4f11f3d26aae96f38084385c322ee3b16fb` | 4,866,413 | BSD-3-Clause |
| LICENSE-Real-ESRGAN.txt | https://raw.githubusercontent.com/xinntao/Real-ESRGAN/a4abfb2979a7bbff3f69f58f58ae324608821e27/LICENSE | `4a699ec4863d96a91fc265948a0c90033f7e8735d515524dcf3444736406e0c2` | 1,519 | BSD-3-Clause |

The existing runtime is MIT. The actual licence heading says “MIT License”; permission includes “use, copy, modify, merge, publish, distribute, sublicense, and/or sell”. Full notice is retained in ../onnxruntime-1.22.0/LICENSE. The reused fflate 0.8.3 MIT notice is retained in ../LICENSE-fflate.txt with the same permission grant. Nothing in these folders was modified. fflate loads only on Download all, using its asynchronous ZIP API and no compression for already-compressed pictures.

| Reused file (relative to this folder) | Source URL | SHA-256 | Bytes | Licence |
| --- | --- | --- | --- | --- |
| ../onnxruntime-1.22.0/ort.wasm.min.js | https://registry.npmjs.org/onnxruntime-web/-/onnxruntime-web-1.22.0.tgz#member=package/dist/ort.wasm.min.js | `65e09376df69107e881b5c34d2d37aed333a366b6d941073ec518168e269b87d` | 48,327 | MIT |
| ../onnxruntime-1.22.0/ort-wasm-simd-threaded.mjs | https://registry.npmjs.org/onnxruntime-web/-/onnxruntime-web-1.22.0.tgz#member=package/dist/ort-wasm-simd-threaded.mjs | `30dd851d9c00622940500f71ddd2ff8820c5cb65270816080175b958705385a8` | 20,856 | MIT |
| ../onnxruntime-1.22.0/ort-wasm-simd-threaded.wasm | https://registry.npmjs.org/onnxruntime-web/-/onnxruntime-web-1.22.0.tgz#member=package/dist/ort-wasm-simd-threaded.wasm | `71aef04959c5c1b6de461b6538e2058e306610034a85aad2742d0c7fd4533fe4` | 11,210,254 | MIT |
| ../onnxruntime-1.22.0/LICENSE | https://raw.githubusercontent.com/microsoft/onnxruntime/v1.22.0/LICENSE | `2f07c72751aed99790b8a4869cf2311df85a860b22ded05fa22803587a48922c` | 1,073 | MIT |
| ../fflate.min.js | https://registry.npmjs.org/fflate/-/fflate-0.8.3.tgz#member=package/umd/index.js | `462ef8041fc970e3615a20a9dd2b2e3047a073b2da729ef4f02b634bba8b7b83` | 33,044 | MIT |
| ../LICENSE-fflate.txt | https://registry.npmjs.org/fflate/-/fflate-0.8.3.tgz#member=package/LICENSE | `0a1df3a083d0c010560aa342e87959c8c1070e6fd54545741f083f22d0c8b551` | 1,069 | MIT |

## Quality decision

The shipped default, Sharp, is the raw Real-ESRGAN RGB model output (clamped and rounded for export). Soft is an optional 50/50 blend of that output and browser Catmull-Rom bicubic; alpha remains plain bicubic for both looks. Real Chromium with ORT 1.22.0 measures all three outputs on the same deterministic test-time PIL fixture:

| Output | PSNR (dB) | SSIM | Mean absolute Laplacian of luma |
| --- | --- | --- | --- |
| Sharp (default, raw model) | 24.8200 | 0.843888 | 5.883315 |
| Soft (optional 50/50 blend) | 26.0437 | 0.865302 | 3.612905 |
| PIL bicubic | 25.2257 | 0.844298 | 1.912643 |
| Original fixture | — | — | 9.054895 |

Sharpness uses the mean absolute four-neighbour Laplacian of Rec.709 luma on the 0–255 scale, excluding the outermost pixel border. The quality gates require Soft PSNR above bicubic, Sharp Laplacian above bicubic, and Sharp PSNR at least bicubic minus 1.5 dB to catch broken inference or colour/offset errors. PSNR against a bicubic-downscaled fixture penalizes sharpening, so it no longer requires raw model output to beat bicubic PSNR. Soft trades sharpening for fidelity on that fixture; users can select it when edges look over-sharpened or faces look painted.

The next candidates were inspected from the actual npm tarballs (not the GitHub sidebar):

| Inspected archive / member (not shipped) | URL | SHA-256 | Bytes |
| --- | --- | --- | --- |
| esrgan-slim 1.0.0 archive | https://registry.npmjs.org/@upscalerjs/esrgan-slim/-/esrgan-slim-1.0.0.tgz | `b53070ff01363e74d33138106dcc72b58d341c9fb7e8fd648554a4e7744142b5` | 4,357,056 |
| esrgan-medium 1.0.0 archive | https://registry.npmjs.org/@upscalerjs/esrgan-medium/-/esrgan-medium-1.0.0.tgz | `367471ded0629387b9aabc75183be0b80d39d46c3f509945ebde4f5bf7a49627` | 11,361,705 |

Both package/LICENSE files quote “Permission is hereby granted, free of charge” and permit redistribution with the notice retained. These package releases actually describe their slim/medium networks as `rdn`, with 0–255 input/output. Their x4 weights were evaluated by exact HWIO-to-OIHW convolution conversion without training: PSNR slim 24.9144 dB and medium 25.0214 dB on the unchanged fixture (torch conversion check). Neither improved the bicubic PSNR baseline in that evaluation. They do not require or justify shipping TensorFlow.js; route B would not change their accuracy.

## Model notice (verbatim)

```text
BSD 3-Clause License

Copyright (c) 2021, Xintao Wang
All rights reserved.

Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:

1. Redistributions of source code must retain the above copyright notice, this
   list of conditions and the following disclaimer.

2. Redistributions in binary form must reproduce the above copyright notice,
   this list of conditions and the following disclaimer in the documentation
   and/or other materials provided with the distribution.

3. Neither the name of the copyright holder nor the names of its
   contributors may be used to endorse or promote products derived from
   this software without specific prior written permission.

THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
```

## Processing and limits

Nothing in /vendor/ loads before Upscale. Worker downloads are same-origin GETs with redirects refused; the model is checked for exact bytes and SHA-256 on download AND cache read. Cache API keys include the model SHA-256. Storage failure still permits inference in memory. Cancel terminates the worker and its active WASM call. Pictures never enter that cache and never leave the browser.

128 input-pixel tile cores keep convolution working memory bounded. Each core extends 8 input pixels for blending and a further 40 pixels for context. The 34 convolutions have a receptive-field radius of 34 pixels; the 40-pixel halo exceeds that radius, so artificial tile padding is discarded. Linear edge ramps are normalized per axis, and products give a partition of unity: every output pixel receives a total weight of one. True image edges keep the official network's zero padding. Alpha uses plain Catmull-Rom bicubic. 2x reduces native 4x with an antialiased bicubic kernel. Both scales cap the intermediate 4x at 4096 pixels on the long edge and 16,000,000 pixels; large inputs require an explicit shrink selection.

Twenty pictures run sequentially, and completed outputs are kept as PNG blobs between exports. Full-resolution blending still needs roughly 256 MB at the largest size plus model workspace, decoding and resampling; phone memory and speed vary. Safari and real phones require device verification. The model sharpens and invents plausible detail; it cannot recover missing information, and faces can look smoothed.
