mirror of
https://github.com/opencv/opencv.git
synced 2026-10-05 04:03:32 +03:00
core: enable CV_64F arithmetic SIMD kernels for scalable SIMD - #30069 ### Summary The CV_64F element-wise arithmetic kernels in `modules/core/src/arithm.simd.hpp` are guarded by `CV_SIMD_64F`, which is 0 on scalable-vector (RVV) builds, while the kernels themselves are already written with width-independent universal intrinsics (`v_float64`, `vx_load`, `VTraits<>::vlanes()`). On RVV every CV_64F arithmetic operation therefore falls back to the scalar kernels even though `CV_SIMD_SCALABLE_64F` is 1 and `v_float64` is fully supported. The same file already uses the dual guard `CV_SIMD_64F || CV_SIMD_SCALABLE_64F` in four other places (see the comment at L166: "scalable (RVV) has v_float64 without CV_SIMD_64F"), so the 12 remaining single-guard sites look like an oversight. This PR widens them. ### What becomes vectorized on RVV (VLEN=256 => 4 double lanes) - `cv::add` / `cv::subtract` (CV_64F) - `cv::multiply` (CV_64F, plus the 32U/32S f64-work-type kernels) - `cv::divide` (DIVW group: 32U/32S/64U/64S/64F) - `cv::min` / `cv::max` (CV_64F) - `cv::absdiff` (CV_64F) - `cv::compare` (CV_64F) - `cv::addWeighted` (AWD group: 32U/32S/64U/64S/64F) All affected kernels are element-wise with no cross-lane reductions, so results are identical to the scalar kernels (per-element IEEE operations). The only theoretical difference is that the addWeighted vector body uses fused `v_fma` while the scalar tail does separate mul+add (at most 1 ulp); the same code already passes x86 AVX2/AVX512 CI where FMA is used as well. On fixed-width backends (x86, AArch64, LoongArch, MIPS MSA, Power VSX) `CV_SIMD_64F` is already 1, so the guard change compiles to exactly the same code as before — no impact on other architectures. ARMv7/AArch32 and WASM SIMD128 have no f64 vector type at all (`CV_SIMD128_64F=0` by design) and keep the scalar path unchanged. ### Scope Generic `rv64gc + CPU_DISPATCH=RVV` builds do not benefit yet (`arithm` is not a dispatched file); the win applies to builds where the core library itself is compiled with RVV enabled (`CPU_BASELINE=RVV`, the common embedded configuration on SpacemiT K1/K3 boards). ### Build configuration Cross-compiled for SpacemiT K1 (8x 1.6 GHz, RVV 1.0, VLEN=256), GCC 14.2.1, `-O3`, both sides built with identical flags: ```bash cmake -GNinja -DCMAKE_TOOLCHAIN_FILE=riscv64-spacemit-k1.toolchain.cmake \ -DCMAKE_BUILD_TYPE=Release -DCPU_BASELINE=RVV -DCPU_DISPATCH= \ -DBUILD_LIST=core,ts,calib,geometry -DBUILD_TESTS=ON -DBUILD_PERF_TESTS=ON \ -DWITH_LAPACK=OFF -DWITH_EIGEN=OFF -DWITH_OPENCL=OFF \ -DWITH_TIFF=OFF -DWITH_OPENJPEG=OFF -DWITH_JASPER=OFF -DWITH_OPENEXR=OFF \ -DWITH_GDAL=OFF -DWITH_GDCM=OFF -DWITH_AVIF=OFF -DWITH_JPEGXL=OFF -DWITH_IMGCODEC_GIF=OFF .. ninja opencv_core opencv_test_core opencv_perf_core opencv_test_calib opencv_test_geometry ``` ### Correctness validation (on board) ```bash export OPENCV_TEST_DATA_PATH=.../opencv_extra/testdata export OPENCV_OPENCL_RUNTIME=disabled ./bin/opencv_test_core --test_threads=4 ./bin/opencv_test_calib --test_threads=4 ./bin/opencv_test_geometry --test_threads=4 ``` | suite | baseline | with this PR | |---|---|---| | `opencv_test_core` (16757 tests) | 16756 pass, 1 fail: `Samples.findFile` (missing samples data path, environment) | **identical failure set** | | `opencv_test_calib` (34 tests) | 34/34 pass | 34/34 pass | | `opencv_test_geometry` (323 tests) | 323/323 pass | 323/323 pass | Zero new failures; the arithm accuracy suites (`Core_*ElemWiseTest`, `Core_AddWeighted*`, `Core_*Mixed/ArithmMixedTest`, compare) exercise the newly vectorized kernels at all depths and pass unchanged. ### Performance (SpacemiT K1, single-threaded) This PR adds a CV_64F perf fixture (`F64ArithmTest`: 9 ops x 4 sizes x CV_64FC1 = 36 cases) to `modules/core/perf/perf_arithm.cpp`, since the existing `BinaryOpTest` fixture does not parametrize CV_64F (and has no `compare`/`addWeighted` cases at all). Numbers below are per-case minima over 4 alternating A/B rounds x 50 samples, reported with the standard `modules/ts/misc/summary.py` (`-m min`). With the default build flags (`-O3`, GCC 14.2.1): ``` Min (ms) Name of Test base cand cand min4 min4 min4 vs base min4 (x-factor) absdiff::F64ArithmTest::(127x61, 64FC1) 0.035 0.026 1.33 absdiff::F64ArithmTest::(640x480, 64FC1) 1.580 1.644 0.96 absdiff::F64ArithmTest::(1280x720, 64FC1) 4.533 5.049 0.90 absdiff::F64ArithmTest::(1920x1080, 64FC1) 10.059 10.936 0.92 add::F64ArithmTest::(127x61, 64FC1) 0.027 0.028 0.98 add::F64ArithmTest::(640x480, 64FC1) 1.529 1.561 0.98 add::F64ArithmTest::(1280x720, 64FC1) 4.560 4.450 1.02 add::F64ArithmTest::(1920x1080, 64FC1) 9.544 9.549 1.00 addWeighted::F64ArithmTest::(127x61, 64FC1) 0.033 0.028 1.14 addWeighted::F64ArithmTest::(640x480, 64FC1) 1.582 1.581 1.00 addWeighted::F64ArithmTest::(1280x720, 64FC1) 4.579 4.797 0.95 addWeighted::F64ArithmTest::(1920x1080, 64FC1) 10.255 9.885 1.04 compare::F64ArithmTest::(127x61, 64FC1) 0.099 0.035 2.85 compare::F64ArithmTest::(640x480, 64FC1) 3.834 1.300 2.95 compare::F64ArithmTest::(1280x720, 64FC1) 11.456 3.847 2.98 compare::F64ArithmTest::(1920x1080, 64FC1) 25.670 8.738 2.94 divide::F64ArithmTest::(127x61, 64FC1) 0.064 0.065 0.98 divide::F64ArithmTest::(640x480, 64FC1) 2.493 2.352 1.06 divide::F64ArithmTest::(1280x720, 64FC1) 7.359 7.192 1.02 divide::F64ArithmTest::(1920x1080, 64FC1) 16.306 15.548 1.05 max::F64ArithmTest::(127x61, 64FC1) 0.029 0.026 1.12 max::F64ArithmTest::(640x480, 64FC1) 1.583 1.600 0.99 max::F64ArithmTest::(1280x720, 64FC1) 4.746 4.963 0.96 max::F64ArithmTest::(1920x1080, 64FC1) 9.827 10.143 0.97 min::F64ArithmTest::(127x61, 64FC1) 0.029 0.026 1.10 min::F64ArithmTest::(640x480, 64FC1) 1.625 1.630 1.00 min::F64ArithmTest::(1280x720, 64FC1) 4.821 5.000 0.96 min::F64ArithmTest::(1920x1080, 64FC1) 9.363 11.126 0.84 multiply::F64ArithmTest::(127x61, 64FC1) 0.029 0.027 1.10 multiply::F64ArithmTest::(640x480, 64FC1) 1.680 1.762 0.95 multiply::F64ArithmTest::(1280x720, 64FC1) 4.835 4.929 0.98 multiply::F64ArithmTest::(1920x1080, 64FC1) 10.427 10.564 0.99 subtract::F64ArithmTest::(127x61, 64FC1) 0.027 0.026 1.06 subtract::F64ArithmTest::(640x480, 64FC1) 1.605 1.660 0.97 subtract::F64ArithmTest::(1280x720, 64FC1) 4.758 5.016 0.95 subtract::F64ArithmTest::(1920x1080, 64FC1) 11.156 10.594 1.05 ``` x-factor geomean **1.13x**. `compare` is a robust ~2.9x; most other ops land near parity because GCC already auto-vectorizes the simple scalar f64 loops at `-O3` on RVV (exactly the issue #30066 works around with `-fno-tree-vectorize`). Per-case values near parity swing with the known K1 bimodal timing behavior. With `-fno-tree-vectorize` added to the build flags (the direction proposed in #30066 for RVV), the same A/B shows the kernel-level gains directly: ``` Min (ms) Name of Test nt nt nt base cand cand min4 min4 min4 vs nt base min4 (x-factor) absdiff::F64ArithmTest::(127x61, 64FC1) 0.087 0.026 3.41 absdiff::F64ArithmTest::(640x480, 64FC1) 3.351 1.423 2.36 absdiff::F64ArithmTest::(1280x720, 64FC1) 10.189 4.595 2.22 absdiff::F64ArithmTest::(1920x1080, 64FC1) 22.476 10.094 2.23 add::F64ArithmTest::(127x61, 64FC1) 0.037 0.027 1.35 add::F64ArithmTest::(640x480, 64FC1) 1.538 1.446 1.06 add::F64ArithmTest::(1280x720, 64FC1) 4.644 4.403 1.05 add::F64ArithmTest::(1920x1080, 64FC1) 10.316 9.629 1.07 addWeighted::F64ArithmTest::(127x61, 64FC1) 0.076 0.027 2.81 addWeighted::F64ArithmTest::(640x480, 64FC1) 2.978 1.363 2.18 addWeighted::F64ArithmTest::(1280x720, 64FC1) 9.141 4.118 2.22 addWeighted::F64ArithmTest::(1920x1080, 64FC1) 19.856 9.075 2.19 compare::F64ArithmTest::(127x61, 64FC1) 0.103 0.031 3.31 compare::F64ArithmTest::(640x480, 64FC1) 3.993 1.225 3.26 compare::F64ArithmTest::(1280x720, 64FC1) 11.801 3.596 3.28 compare::F64ArithmTest::(1920x1080, 64FC1) 26.611 8.125 3.28 divide::F64ArithmTest::(127x61, 64FC1) 0.156 0.065 2.41 divide::F64ArithmTest::(640x480, 64FC1) 6.095 2.330 2.62 divide::F64ArithmTest::(1280x720, 64FC1) 18.191 6.880 2.64 divide::F64ArithmTest::(1920x1080, 64FC1) 40.743 15.331 2.66 max::F64ArithmTest::(127x61, 64FC1) 0.085 0.025 3.36 max::F64ArithmTest::(640x480, 64FC1) 3.272 1.428 2.29 max::F64ArithmTest::(1280x720, 64FC1) 10.009 4.581 2.18 max::F64ArithmTest::(1920x1080, 64FC1) 21.962 10.130 2.17 min::F64ArithmTest::(127x61, 64FC1) 0.086 0.025 3.40 min::F64ArithmTest::(640x480, 64FC1) 3.316 1.417 2.34 min::F64ArithmTest::(1280x720, 64FC1) 10.133 4.583 2.21 min::F64ArithmTest::(1920x1080, 64FC1) 22.334 10.160 2.20 multiply::F64ArithmTest::(127x61, 64FC1) 0.062 0.024 2.61 multiply::F64ArithmTest::(640x480, 64FC1) 2.430 1.362 1.78 multiply::F64ArithmTest::(1280x720, 64FC1) 7.427 4.344 1.71 multiply::F64ArithmTest::(1920x1080, 64FC1) 16.160 9.370 1.72 subtract::F64ArithmTest::(127x61, 64FC1) 0.038 0.025 1.48 subtract::F64ArithmTest::(640x480, 64FC1) 1.696 1.523 1.11 subtract::F64ArithmTest::(1280x720, 64FC1) 4.794 4.805 1.00 subtract::F64ArithmTest::(1920x1080, 64FC1) 12.058 10.820 1.11 ``` x-factor geomean **2.09x** (36 cases, 1.00-3.41x): compare ~3.3x, divide ~2.6x, min/max ~2.2x, absdiff ~2.2x, addWeighted ~2.2x, multiply ~1.7-2.6x. `add`/`subtract` stay near parity in both regimes — they appear to be served by a different (already-vectorized) path before the kernel table is consulted. ### Pull Request Readiness Checklist See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request - [x] I agree to contribute to the project under Apache 2 License. - [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV - [x] The PR is proposed to the proper branch - [x] There is a reference to the original bug report and related work - [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable Patch to opencv_extra has the same branch name. - [ ] The feature is well documented and sample code can be built with the project CMake