Files
opencv/modules/core
Teddy-Yangjiale 95e8141250 Merge pull request #30069 from Teddy-Yangjiale:rvv-f64-arithm
core: enable CV_64F arithmetic SIMD kernels for scalable SIMD - #30069

### Summary

The CV_64F element-wise arithmetic kernels in `modules/core/src/arithm.simd.hpp` are guarded by `CV_SIMD_64F`, which is 0 on scalable-vector (RVV) builds, while the kernels themselves are already written with width-independent universal intrinsics (`v_float64`, `vx_load`, `VTraits<>::vlanes()`). On RVV every CV_64F arithmetic operation therefore falls back to the scalar kernels even though `CV_SIMD_SCALABLE_64F` is 1 and `v_float64` is fully supported.

The same file already uses the dual guard `CV_SIMD_64F || CV_SIMD_SCALABLE_64F` in four other places (see the comment at L166: "scalable (RVV) has v_float64 without CV_SIMD_64F"), so the 12 remaining single-guard sites look like an oversight. This PR widens them.

### What becomes vectorized on RVV (VLEN=256 => 4 double lanes)

- `cv::add` / `cv::subtract` (CV_64F)
- `cv::multiply` (CV_64F, plus the 32U/32S f64-work-type kernels)
- `cv::divide` (DIVW group: 32U/32S/64U/64S/64F)
- `cv::min` / `cv::max` (CV_64F)
- `cv::absdiff` (CV_64F)
- `cv::compare` (CV_64F)
- `cv::addWeighted` (AWD group: 32U/32S/64U/64S/64F)

All affected kernels are element-wise with no cross-lane reductions, so results are identical to the scalar kernels (per-element IEEE operations). The only theoretical difference is that the addWeighted vector body uses fused `v_fma` while the scalar tail does separate mul+add (at most 1 ulp); the same code already passes x86 AVX2/AVX512 CI where FMA is used as well.

On fixed-width backends (x86, AArch64, LoongArch, MIPS MSA, Power VSX) `CV_SIMD_64F` is already 1, so the guard change compiles to exactly the same code as before — no impact on other architectures. ARMv7/AArch32 and WASM SIMD128 have no f64 vector type at all (`CV_SIMD128_64F=0` by design) and keep the scalar path unchanged.

### Scope

Generic `rv64gc + CPU_DISPATCH=RVV` builds do not benefit yet (`arithm` is not a dispatched file); the win applies to builds where the core library itself is compiled with RVV enabled (`CPU_BASELINE=RVV`, the common embedded configuration on SpacemiT K1/K3 boards).

### Build configuration

Cross-compiled for SpacemiT K1 (8x 1.6 GHz, RVV 1.0, VLEN=256), GCC 14.2.1, `-O3`, both sides built with identical flags:

```bash
cmake -GNinja -DCMAKE_TOOLCHAIN_FILE=riscv64-spacemit-k1.toolchain.cmake \
      -DCMAKE_BUILD_TYPE=Release -DCPU_BASELINE=RVV -DCPU_DISPATCH= \
      -DBUILD_LIST=core,ts,calib,geometry -DBUILD_TESTS=ON -DBUILD_PERF_TESTS=ON \
      -DWITH_LAPACK=OFF -DWITH_EIGEN=OFF -DWITH_OPENCL=OFF \
      -DWITH_TIFF=OFF -DWITH_OPENJPEG=OFF -DWITH_JASPER=OFF -DWITH_OPENEXR=OFF \
      -DWITH_GDAL=OFF -DWITH_GDCM=OFF -DWITH_AVIF=OFF -DWITH_JPEGXL=OFF -DWITH_IMGCODEC_GIF=OFF ..

ninja opencv_core opencv_test_core opencv_perf_core opencv_test_calib opencv_test_geometry
```


### Correctness validation (on board)

```bash
export OPENCV_TEST_DATA_PATH=.../opencv_extra/testdata
export OPENCV_OPENCL_RUNTIME=disabled
./bin/opencv_test_core     --test_threads=4
./bin/opencv_test_calib    --test_threads=4
./bin/opencv_test_geometry --test_threads=4
```

| suite | baseline | with this PR |
|---|---|---|
| `opencv_test_core` (16757 tests) | 16756 pass, 1 fail: `Samples.findFile` (missing samples data path, environment) | **identical failure set** |
| `opencv_test_calib` (34 tests) | 34/34 pass | 34/34 pass |
| `opencv_test_geometry` (323 tests) | 323/323 pass | 323/323 pass |

Zero new failures; the arithm accuracy suites (`Core_*ElemWiseTest`, `Core_AddWeighted*`, `Core_*Mixed/ArithmMixedTest`, compare) exercise the newly vectorized kernels at all depths and pass unchanged.

### Performance (SpacemiT K1, single-threaded)

This PR adds a CV_64F perf fixture (`F64ArithmTest`: 9 ops x 4 sizes x CV_64FC1 = 36 cases) to `modules/core/perf/perf_arithm.cpp`, since the existing `BinaryOpTest` fixture does not parametrize CV_64F (and has no `compare`/`addWeighted` cases at all). Numbers below are per-case minima over 4 alternating A/B rounds x 50 samples, reported with the standard `modules/ts/misc/summary.py` (`-m min`).

With the default build flags (`-O3`, GCC 14.2.1):

```
Min (ms)

                 Name of Test                   base   cand     cand
                                                min4   min4     min4
                                                                 vs
                                                                base
                                                                min4
                                                             (x-factor)
absdiff::F64ArithmTest::(127x61, 64FC1)        0.035  0.026     1.33
absdiff::F64ArithmTest::(640x480, 64FC1)       1.580  1.644     0.96
absdiff::F64ArithmTest::(1280x720, 64FC1)      4.533  5.049     0.90
absdiff::F64ArithmTest::(1920x1080, 64FC1)     10.059 10.936    0.92
add::F64ArithmTest::(127x61, 64FC1)            0.027  0.028     0.98
add::F64ArithmTest::(640x480, 64FC1)           1.529  1.561     0.98
add::F64ArithmTest::(1280x720, 64FC1)          4.560  4.450     1.02
add::F64ArithmTest::(1920x1080, 64FC1)         9.544  9.549     1.00
addWeighted::F64ArithmTest::(127x61, 64FC1)    0.033  0.028     1.14
addWeighted::F64ArithmTest::(640x480, 64FC1)   1.582  1.581     1.00
addWeighted::F64ArithmTest::(1280x720, 64FC1)  4.579  4.797     0.95
addWeighted::F64ArithmTest::(1920x1080, 64FC1) 10.255 9.885     1.04
compare::F64ArithmTest::(127x61, 64FC1)        0.099  0.035     2.85
compare::F64ArithmTest::(640x480, 64FC1)       3.834  1.300     2.95
compare::F64ArithmTest::(1280x720, 64FC1)      11.456 3.847     2.98
compare::F64ArithmTest::(1920x1080, 64FC1)     25.670 8.738     2.94
divide::F64ArithmTest::(127x61, 64FC1)         0.064  0.065     0.98
divide::F64ArithmTest::(640x480, 64FC1)        2.493  2.352     1.06
divide::F64ArithmTest::(1280x720, 64FC1)       7.359  7.192     1.02
divide::F64ArithmTest::(1920x1080, 64FC1)      16.306 15.548    1.05
max::F64ArithmTest::(127x61, 64FC1)            0.029  0.026     1.12
max::F64ArithmTest::(640x480, 64FC1)           1.583  1.600     0.99
max::F64ArithmTest::(1280x720, 64FC1)          4.746  4.963     0.96
max::F64ArithmTest::(1920x1080, 64FC1)         9.827  10.143    0.97
min::F64ArithmTest::(127x61, 64FC1)            0.029  0.026     1.10
min::F64ArithmTest::(640x480, 64FC1)           1.625  1.630     1.00
min::F64ArithmTest::(1280x720, 64FC1)          4.821  5.000     0.96
min::F64ArithmTest::(1920x1080, 64FC1)         9.363  11.126    0.84
multiply::F64ArithmTest::(127x61, 64FC1)       0.029  0.027     1.10
multiply::F64ArithmTest::(640x480, 64FC1)      1.680  1.762     0.95
multiply::F64ArithmTest::(1280x720, 64FC1)     4.835  4.929     0.98
multiply::F64ArithmTest::(1920x1080, 64FC1)    10.427 10.564    0.99
subtract::F64ArithmTest::(127x61, 64FC1)       0.027  0.026     1.06
subtract::F64ArithmTest::(640x480, 64FC1)      1.605  1.660     0.97
subtract::F64ArithmTest::(1280x720, 64FC1)     4.758  5.016     0.95
subtract::F64ArithmTest::(1920x1080, 64FC1)    11.156 10.594    1.05
```

x-factor geomean **1.13x**. `compare` is a robust ~2.9x; most other ops land near parity because GCC already auto-vectorizes the simple scalar f64 loops at `-O3` on RVV (exactly the issue #30066 works around with `-fno-tree-vectorize`). Per-case values near parity swing with the known K1 bimodal timing behavior.

With `-fno-tree-vectorize` added to the build flags (the direction proposed in #30066 for RVV), the same A/B shows the kernel-level gains directly:

```
Min (ms)

                 Name of Test                    nt     nt       nt
                                                base   cand     cand
                                                min4   min4     min4
                                                                 vs
                                                                 nt
                                                                base
                                                                min4
                                                             (x-factor)
absdiff::F64ArithmTest::(127x61, 64FC1)        0.087  0.026     3.41
absdiff::F64ArithmTest::(640x480, 64FC1)       3.351  1.423     2.36
absdiff::F64ArithmTest::(1280x720, 64FC1)      10.189 4.595     2.22
absdiff::F64ArithmTest::(1920x1080, 64FC1)     22.476 10.094    2.23
add::F64ArithmTest::(127x61, 64FC1)            0.037  0.027     1.35
add::F64ArithmTest::(640x480, 64FC1)           1.538  1.446     1.06
add::F64ArithmTest::(1280x720, 64FC1)          4.644  4.403     1.05
add::F64ArithmTest::(1920x1080, 64FC1)         10.316 9.629     1.07
addWeighted::F64ArithmTest::(127x61, 64FC1)    0.076  0.027     2.81
addWeighted::F64ArithmTest::(640x480, 64FC1)   2.978  1.363     2.18
addWeighted::F64ArithmTest::(1280x720, 64FC1)  9.141  4.118     2.22
addWeighted::F64ArithmTest::(1920x1080, 64FC1) 19.856 9.075     2.19
compare::F64ArithmTest::(127x61, 64FC1)        0.103  0.031     3.31
compare::F64ArithmTest::(640x480, 64FC1)       3.993  1.225     3.26
compare::F64ArithmTest::(1280x720, 64FC1)      11.801 3.596     3.28
compare::F64ArithmTest::(1920x1080, 64FC1)     26.611 8.125     3.28
divide::F64ArithmTest::(127x61, 64FC1)         0.156  0.065     2.41
divide::F64ArithmTest::(640x480, 64FC1)        6.095  2.330     2.62
divide::F64ArithmTest::(1280x720, 64FC1)       18.191 6.880     2.64
divide::F64ArithmTest::(1920x1080, 64FC1)      40.743 15.331    2.66
max::F64ArithmTest::(127x61, 64FC1)            0.085  0.025     3.36
max::F64ArithmTest::(640x480, 64FC1)           3.272  1.428     2.29
max::F64ArithmTest::(1280x720, 64FC1)          10.009 4.581     2.18
max::F64ArithmTest::(1920x1080, 64FC1)         21.962 10.130    2.17
min::F64ArithmTest::(127x61, 64FC1)            0.086  0.025     3.40
min::F64ArithmTest::(640x480, 64FC1)           3.316  1.417     2.34
min::F64ArithmTest::(1280x720, 64FC1)          10.133 4.583     2.21
min::F64ArithmTest::(1920x1080, 64FC1)         22.334 10.160    2.20
multiply::F64ArithmTest::(127x61, 64FC1)       0.062  0.024     2.61
multiply::F64ArithmTest::(640x480, 64FC1)      2.430  1.362     1.78
multiply::F64ArithmTest::(1280x720, 64FC1)     7.427  4.344     1.71
multiply::F64ArithmTest::(1920x1080, 64FC1)    16.160 9.370     1.72
subtract::F64ArithmTest::(127x61, 64FC1)       0.038  0.025     1.48
subtract::F64ArithmTest::(640x480, 64FC1)      1.696  1.523     1.11
subtract::F64ArithmTest::(1280x720, 64FC1)     4.794  4.805     1.00
subtract::F64ArithmTest::(1920x1080, 64FC1)    12.058 10.820    1.11
```

x-factor geomean **2.09x** (36 cases, 1.00-3.41x): compare ~3.3x, divide ~2.6x, min/max ~2.2x, absdiff ~2.2x, addWeighted ~2.2x, multiply ~1.7-2.6x. `add`/`subtract` stay near parity in both regimes — they appear to be served by a different (already-vectorized) path before the kernel table is consulted.



### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
2026-09-29 08:31:04 +03:00
..
2026-05-25 17:49:25 +03:00