core: enable CV_64F arithmetic SIMD kernels for scalable SIMD - #30069
### Summary
The CV_64F element-wise arithmetic kernels in `modules/core/src/arithm.simd.hpp` are guarded by `CV_SIMD_64F`, which is 0 on scalable-vector (RVV) builds, while the kernels themselves are already written with width-independent universal intrinsics (`v_float64`, `vx_load`, `VTraits<>::vlanes()`). On RVV every CV_64F arithmetic operation therefore falls back to the scalar kernels even though `CV_SIMD_SCALABLE_64F` is 1 and `v_float64` is fully supported.
The same file already uses the dual guard `CV_SIMD_64F || CV_SIMD_SCALABLE_64F` in four other places (see the comment at L166: "scalable (RVV) has v_float64 without CV_SIMD_64F"), so the 12 remaining single-guard sites look like an oversight. This PR widens them.
### What becomes vectorized on RVV (VLEN=256 => 4 double lanes)
- `cv::add` / `cv::subtract` (CV_64F)
- `cv::multiply` (CV_64F, plus the 32U/32S f64-work-type kernels)
- `cv::divide` (DIVW group: 32U/32S/64U/64S/64F)
- `cv::min` / `cv::max` (CV_64F)
- `cv::absdiff` (CV_64F)
- `cv::compare` (CV_64F)
- `cv::addWeighted` (AWD group: 32U/32S/64U/64S/64F)
All affected kernels are element-wise with no cross-lane reductions, so results are identical to the scalar kernels (per-element IEEE operations). The only theoretical difference is that the addWeighted vector body uses fused `v_fma` while the scalar tail does separate mul+add (at most 1 ulp); the same code already passes x86 AVX2/AVX512 CI where FMA is used as well.
On fixed-width backends (x86, AArch64, LoongArch, MIPS MSA, Power VSX) `CV_SIMD_64F` is already 1, so the guard change compiles to exactly the same code as before — no impact on other architectures. ARMv7/AArch32 and WASM SIMD128 have no f64 vector type at all (`CV_SIMD128_64F=0` by design) and keep the scalar path unchanged.
### Scope
Generic `rv64gc + CPU_DISPATCH=RVV` builds do not benefit yet (`arithm` is not a dispatched file); the win applies to builds where the core library itself is compiled with RVV enabled (`CPU_BASELINE=RVV`, the common embedded configuration on SpacemiT K1/K3 boards).
### Build configuration
Cross-compiled for SpacemiT K1 (8x 1.6 GHz, RVV 1.0, VLEN=256), GCC 14.2.1, `-O3`, both sides built with identical flags:
```bash
cmake -GNinja -DCMAKE_TOOLCHAIN_FILE=riscv64-spacemit-k1.toolchain.cmake \
-DCMAKE_BUILD_TYPE=Release -DCPU_BASELINE=RVV -DCPU_DISPATCH= \
-DBUILD_LIST=core,ts,calib,geometry -DBUILD_TESTS=ON -DBUILD_PERF_TESTS=ON \
-DWITH_LAPACK=OFF -DWITH_EIGEN=OFF -DWITH_OPENCL=OFF \
-DWITH_TIFF=OFF -DWITH_OPENJPEG=OFF -DWITH_JASPER=OFF -DWITH_OPENEXR=OFF \
-DWITH_GDAL=OFF -DWITH_GDCM=OFF -DWITH_AVIF=OFF -DWITH_JPEGXL=OFF -DWITH_IMGCODEC_GIF=OFF ..
ninja opencv_core opencv_test_core opencv_perf_core opencv_test_calib opencv_test_geometry
```
### Correctness validation (on board)
```bash
export OPENCV_TEST_DATA_PATH=.../opencv_extra/testdata
export OPENCV_OPENCL_RUNTIME=disabled
./bin/opencv_test_core --test_threads=4
./bin/opencv_test_calib --test_threads=4
./bin/opencv_test_geometry --test_threads=4
```
| suite | baseline | with this PR |
|---|---|---|
| `opencv_test_core` (16757 tests) | 16756 pass, 1 fail: `Samples.findFile` (missing samples data path, environment) | **identical failure set** |
| `opencv_test_calib` (34 tests) | 34/34 pass | 34/34 pass |
| `opencv_test_geometry` (323 tests) | 323/323 pass | 323/323 pass |
Zero new failures; the arithm accuracy suites (`Core_*ElemWiseTest`, `Core_AddWeighted*`, `Core_*Mixed/ArithmMixedTest`, compare) exercise the newly vectorized kernels at all depths and pass unchanged.
### Performance (SpacemiT K1, single-threaded)
This PR adds a CV_64F perf fixture (`F64ArithmTest`: 9 ops x 4 sizes x CV_64FC1 = 36 cases) to `modules/core/perf/perf_arithm.cpp`, since the existing `BinaryOpTest` fixture does not parametrize CV_64F (and has no `compare`/`addWeighted` cases at all). Numbers below are per-case minima over 4 alternating A/B rounds x 50 samples, reported with the standard `modules/ts/misc/summary.py` (`-m min`).
With the default build flags (`-O3`, GCC 14.2.1):
```
Min (ms)
Name of Test base cand cand
min4 min4 min4
vs
base
min4
(x-factor)
absdiff::F64ArithmTest::(127x61, 64FC1) 0.035 0.026 1.33
absdiff::F64ArithmTest::(640x480, 64FC1) 1.580 1.644 0.96
absdiff::F64ArithmTest::(1280x720, 64FC1) 4.533 5.049 0.90
absdiff::F64ArithmTest::(1920x1080, 64FC1) 10.059 10.936 0.92
add::F64ArithmTest::(127x61, 64FC1) 0.027 0.028 0.98
add::F64ArithmTest::(640x480, 64FC1) 1.529 1.561 0.98
add::F64ArithmTest::(1280x720, 64FC1) 4.560 4.450 1.02
add::F64ArithmTest::(1920x1080, 64FC1) 9.544 9.549 1.00
addWeighted::F64ArithmTest::(127x61, 64FC1) 0.033 0.028 1.14
addWeighted::F64ArithmTest::(640x480, 64FC1) 1.582 1.581 1.00
addWeighted::F64ArithmTest::(1280x720, 64FC1) 4.579 4.797 0.95
addWeighted::F64ArithmTest::(1920x1080, 64FC1) 10.255 9.885 1.04
compare::F64ArithmTest::(127x61, 64FC1) 0.099 0.035 2.85
compare::F64ArithmTest::(640x480, 64FC1) 3.834 1.300 2.95
compare::F64ArithmTest::(1280x720, 64FC1) 11.456 3.847 2.98
compare::F64ArithmTest::(1920x1080, 64FC1) 25.670 8.738 2.94
divide::F64ArithmTest::(127x61, 64FC1) 0.064 0.065 0.98
divide::F64ArithmTest::(640x480, 64FC1) 2.493 2.352 1.06
divide::F64ArithmTest::(1280x720, 64FC1) 7.359 7.192 1.02
divide::F64ArithmTest::(1920x1080, 64FC1) 16.306 15.548 1.05
max::F64ArithmTest::(127x61, 64FC1) 0.029 0.026 1.12
max::F64ArithmTest::(640x480, 64FC1) 1.583 1.600 0.99
max::F64ArithmTest::(1280x720, 64FC1) 4.746 4.963 0.96
max::F64ArithmTest::(1920x1080, 64FC1) 9.827 10.143 0.97
min::F64ArithmTest::(127x61, 64FC1) 0.029 0.026 1.10
min::F64ArithmTest::(640x480, 64FC1) 1.625 1.630 1.00
min::F64ArithmTest::(1280x720, 64FC1) 4.821 5.000 0.96
min::F64ArithmTest::(1920x1080, 64FC1) 9.363 11.126 0.84
multiply::F64ArithmTest::(127x61, 64FC1) 0.029 0.027 1.10
multiply::F64ArithmTest::(640x480, 64FC1) 1.680 1.762 0.95
multiply::F64ArithmTest::(1280x720, 64FC1) 4.835 4.929 0.98
multiply::F64ArithmTest::(1920x1080, 64FC1) 10.427 10.564 0.99
subtract::F64ArithmTest::(127x61, 64FC1) 0.027 0.026 1.06
subtract::F64ArithmTest::(640x480, 64FC1) 1.605 1.660 0.97
subtract::F64ArithmTest::(1280x720, 64FC1) 4.758 5.016 0.95
subtract::F64ArithmTest::(1920x1080, 64FC1) 11.156 10.594 1.05
```
x-factor geomean **1.13x**. `compare` is a robust ~2.9x; most other ops land near parity because GCC already auto-vectorizes the simple scalar f64 loops at `-O3` on RVV (exactly the issue #30066 works around with `-fno-tree-vectorize`). Per-case values near parity swing with the known K1 bimodal timing behavior.
With `-fno-tree-vectorize` added to the build flags (the direction proposed in #30066 for RVV), the same A/B shows the kernel-level gains directly:
```
Min (ms)
Name of Test nt nt nt
base cand cand
min4 min4 min4
vs
nt
base
min4
(x-factor)
absdiff::F64ArithmTest::(127x61, 64FC1) 0.087 0.026 3.41
absdiff::F64ArithmTest::(640x480, 64FC1) 3.351 1.423 2.36
absdiff::F64ArithmTest::(1280x720, 64FC1) 10.189 4.595 2.22
absdiff::F64ArithmTest::(1920x1080, 64FC1) 22.476 10.094 2.23
add::F64ArithmTest::(127x61, 64FC1) 0.037 0.027 1.35
add::F64ArithmTest::(640x480, 64FC1) 1.538 1.446 1.06
add::F64ArithmTest::(1280x720, 64FC1) 4.644 4.403 1.05
add::F64ArithmTest::(1920x1080, 64FC1) 10.316 9.629 1.07
addWeighted::F64ArithmTest::(127x61, 64FC1) 0.076 0.027 2.81
addWeighted::F64ArithmTest::(640x480, 64FC1) 2.978 1.363 2.18
addWeighted::F64ArithmTest::(1280x720, 64FC1) 9.141 4.118 2.22
addWeighted::F64ArithmTest::(1920x1080, 64FC1) 19.856 9.075 2.19
compare::F64ArithmTest::(127x61, 64FC1) 0.103 0.031 3.31
compare::F64ArithmTest::(640x480, 64FC1) 3.993 1.225 3.26
compare::F64ArithmTest::(1280x720, 64FC1) 11.801 3.596 3.28
compare::F64ArithmTest::(1920x1080, 64FC1) 26.611 8.125 3.28
divide::F64ArithmTest::(127x61, 64FC1) 0.156 0.065 2.41
divide::F64ArithmTest::(640x480, 64FC1) 6.095 2.330 2.62
divide::F64ArithmTest::(1280x720, 64FC1) 18.191 6.880 2.64
divide::F64ArithmTest::(1920x1080, 64FC1) 40.743 15.331 2.66
max::F64ArithmTest::(127x61, 64FC1) 0.085 0.025 3.36
max::F64ArithmTest::(640x480, 64FC1) 3.272 1.428 2.29
max::F64ArithmTest::(1280x720, 64FC1) 10.009 4.581 2.18
max::F64ArithmTest::(1920x1080, 64FC1) 21.962 10.130 2.17
min::F64ArithmTest::(127x61, 64FC1) 0.086 0.025 3.40
min::F64ArithmTest::(640x480, 64FC1) 3.316 1.417 2.34
min::F64ArithmTest::(1280x720, 64FC1) 10.133 4.583 2.21
min::F64ArithmTest::(1920x1080, 64FC1) 22.334 10.160 2.20
multiply::F64ArithmTest::(127x61, 64FC1) 0.062 0.024 2.61
multiply::F64ArithmTest::(640x480, 64FC1) 2.430 1.362 1.78
multiply::F64ArithmTest::(1280x720, 64FC1) 7.427 4.344 1.71
multiply::F64ArithmTest::(1920x1080, 64FC1) 16.160 9.370 1.72
subtract::F64ArithmTest::(127x61, 64FC1) 0.038 0.025 1.48
subtract::F64ArithmTest::(640x480, 64FC1) 1.696 1.523 1.11
subtract::F64ArithmTest::(1280x720, 64FC1) 4.794 4.805 1.00
subtract::F64ArithmTest::(1920x1080, 64FC1) 12.058 10.820 1.11
```
x-factor geomean **2.09x** (36 cases, 1.00-3.41x): compare ~3.3x, divide ~2.6x, min/max ~2.2x, absdiff ~2.2x, addWeighted ~2.2x, multiply ~1.7-2.6x. `add`/`subtract` stay near parity in both regimes — they appear to be served by a different (already-vectorized) path before the kernel table is consulted.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
core: RVV HAL transpose2d for 3-channel element sizes - #30075
### Summary
`cv::transpose` for 3-channel types (`8UC3`/`8SC3`, `16UC3`/`16SC3`/`16FC3`, `32SC3`/`32FC3`) falls back to the scalar core path on RISC-V. The RVV HAL `transpose2d` dispatches on *element size* and only implements `esz ∈ {1, 2, 4, 8}`. This PR adds kernels for
`esz = 3, 6, 12`, and an accuracy test that covers the new paths including non-continuous ROI.
### Root cause
- `hal/riscv-rvv/src/core/transpose.cpp`: the `Transpose2dFunc tab[]` table has entries only at indices 1/2/4/8, so `cv_hal_transpose2d` returns `CV_HAL_ERROR_NOT_IMPLEMENTED` for every 3-channel type and the caller runs the generic path.
- On the scalable RVV backend (`intrin_rvv_scalable.hpp`) `CV_SIMD128` is never defined, so the existing `transpose_8/16/32/48bit_simd` fast paths `modules/core/src/matrix_transform.cpp` are compiled out entirely; what runs is the scalar 4x4 element loop. 8UC3 (the default color type) is the most visible case.
### Implementation
The new kernels deinterleave the three channels of several source rows with `vlseg3e{L}` and emit the transposed rows with strided segment stores (`vssseg{6,8}e{L}`). The number of source rows per block was picked experimentally: 8 rows for 8-bit lanes, 2 rows for 16/32-bit lanes. For 16/32-bit lanes, taller blocks e.g. 8 rows will add register pressure and cause regression. A one-row tail covers the remainder; unaligned 16/32-bit inputs fall back to the generic path.
### Performance
SpacemiT K1 / X60, OpenCV 5.x, `--perf_force_samples=20
--perf_min_samples=20`, `BinaryOpTest.transpose2d`:
| type (esz) | 640x480 | 1280x720 | 1920x1080 |
|---|---:|---:|---:|
| `CV_8UC3` (3) | 2.758 -> 1.648 ms (**1.67x**) | 16.772 -> 10.575 ms (**1.59x**) | 42.901 -> 39.552 ms (1.09x) |
| `CV_16SC3` (6) | 8.384 -> 3.556 ms (**2.36x**) | 31.112 -> 19.450 ms (**1.60x**) | 71.921 -> 62.460 ms (1.15x) |
Untouched element sizes (esz 1/2/4/8, 16 cases at 640x480 + 1280x720): geomean **1.005x**,
range 0.977x ... 1.058x.
### Accuracy
New test `Core_Transpose.C3ElementSizesWithRoi` (`modules/core/test/test_mat.cpp`) sweeps
`CV_8UC3, CV_8UC(6), CV_16SC3, CV_16FC3, CV_8UC(12), CV_32FC3` over sizes `1x1, 2x3, 137x5, 133x4` through non-continuous ROI views, byte-compares every element against the source, and asserts the bytes outside the destination ROI are untouched. On K1:
```
[ PASSED ] 4 tests. # Core_Transpose.C3ElementSizesWithRoi
# Core_Transpose/ElemWiseTest.accuracy/0
# Core_Rotate/ElemWiseTest.accuracy/0
```
`BinaryOpTest.transpose2d` already provides the performance coverage; no `opencv_extra` data is needed.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Verify assembler can encode AVX-VNNI before enabling it - #29844
Closes: https://github.com/opencv/opencv/issues/29840
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
v_select had no 64-bit integer form on any backend except WASM, so code
that needs to blend v_int64/v_uint64 through a comparison mask had to
apply the mask by hand. intrin_sse.hpp carried the two entries commented
out as TBD; the other backends simply stopped at 32-bit and float64.
Every backend already blends bitwise or through a byte/boolean cast, so
the lane width does not change the operation:
SSE _mm_blendv_pd under CV_SSE4_1, the xor/and/xor form otherwise
AVX _mm256_blendv_epi8, as the narrower types already use
NEON vbslq_u64 / vbslq_s64
MSA msa_bslq_u8 over the byte reinterpretation
VSX vec_sel with the same boolean cast v_float64x2 uses
LSX/LASX __lsx_vbitsel_v / __lasx_xvbitsel_v on the raw register
RVV 0.7.1 vmerge_vvm with the b64 mask
RVV __riscv_vmerge, as for the narrower types
WASM already had both and is unchanged, and the generic v_reg
implementation in intrin_cpp.hpp already covers every type.
The 64-bit test chains could not simply call test_mask(), because its
v_signmask() expectations assume more than two lanes. Added test_select(),
which exercises v_select alone, and called it from TheTest<v_uint64>() and
TheTest<v_int64>().
Fix function pointer signature mismatches - #30040
Contrib PR: https://github.com/opencv/opencv_contrib/pull/4224
Calling a function through a function pointer with a mismatched signature is undefined behavior in C/C++ and causes Clang Control Flow Integrity to trap with `SIGILL` (`ud1`) at indirect call sites.
This is a follow-up on https://github.com/opencv/opencv/pull/28939
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
readdir() allocates a new buffer for dir->ent.d_name on every call
and overwrites the previous pointer without freeing it. Since
cv::glob() calls readdir() once per directory entry, every call
except the last leaks its allocation. Under _WIN32_WCE, DIR has no
destructor at all, so every allocation leaks, including the last
one.
(cherry picked from commit 20b9529bc3)
Process masked integer norms chunk by chunk and reuse each loaded mask
and predicate across channels, avoiding repeated mask processing.
Extend norm_mask performance coverage with three-channel 8U, 8S, 16U,
16S, and 32S inputs for INF, L1, and L2 norms.
Co-authored-by: Yang Wang <yangwang@iscas.ac.cn>
Co-authored-by: Yuansheng <yuansheng@isrc.iscas.ac.cn>
doc: document the bit layout of Mat::type() - #29981
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
- [ ] The feature is well documented and sample code can be built with the project CMake
### Description
Fixes#24901.
#### Problem
`cv::Mat::type()` returns a packed bit-field, but the encoding was not documented. Users
had to reverse-engineer the layout from the `CV_*` macros to answer questions such as how many
channels fit, how to build a type from a depth and channel count, and which bits are reserved for
the matrix flags.
#### Change
Expanded the Doxygen for `Mat::type()` in `modules/core/include/opencv2/core/mat.hpp` to describe
the layout used by the 5.x branch:
- bits 0-4 (`CV_MAT_DEPTH_MASK`) – element depth (5 bits);
- bits 5-11 (`CV_MAT_CN_MASK`) – number of channels minus one (7 bits), i.e. 1..`CV_CN_MAX` (128);
- together these occupy the lowest 12 bits (`CV_MAT_TYPE_MASK`).
The description also points to `CV_MAT_DEPTH()`, `CV_MAT_CN()` and `CV_MAKETYPE()`, and notes that
the continuity (`CV_MAT_CONT_FLAG`) and submatrix (`CV_SUBMAT_FLAG`) bits of `Mat::flags` are not
part of the returned value.
#### Branch note
This PR targets `5.x` only. The encoding changed between branches: 5.x uses `CV_CN_SHIFT == 5`
(5-bit depth, 7-bit channel count), whereas 4.x uses `CV_CN_SHIFT == 3`. As requested in the issue,
a separate `4.x` PR would be needed for that branch.
#### Verification
Documentation-only change; no code or behavior is modified. The bit ranges and macro names were
checked against `modules/core/include/opencv2/core/hal/interface.h` and
`modules/core/include/opencv2/core/cvdef.h`:
```
CV_CN_MAX 128
CV_CN_SHIFT 5
CV_DEPTH_MAX (1 << CV_CN_SHIFT) // 32
CV_MAT_DEPTH_MASK (CV_DEPTH_MAX - 1) // 0x1F -> bits 0-4
CV_MAT_CN_MASK ((CV_CN_MAX - 1) << 5) // 0xFE0 -> bits 5-11
CV_MAT_TYPE_MASK (CV_DEPTH_MAX*CV_CN_MAX-1) // 0xFFF
```
Extend the CV_SIMD guard in lapack.cpp to include CV_SIMD_SCALABLE,
addressing the TODO updated in PR24325.
Co-authored-by: Yang Wang <yangwang@iscas.ac.cn>
Co-authored-by: Yuansheng <yuansheng@isrc.iscas.ac.cn>
core: optimize float REDUCE_SUM2 with RVV - #29930
### Summary
Add an RVV 1.0 kernel for `cv::reduce` with `dim=0` and `CV_32F`
input/output through the CPU dispatch mechanism.
The optimized path targets the floating-point `REDUCE_SUM2` operation.
### Implementation
Process four source rows per vector iteration to reduce intermediate
buffer traffic while preserving source-row accumulation order.
The RVV kernel uses LMUL=4 and dynamic vector lengths for tail handling.
Output writes are deferred until all input rows have been read. This
preserves correct behavior when the source and destination matrices
overlap.
The implementation is VLEN-agnostic.
### Functional Testing
The Reduce tests were run with:
```bash
./bin/opencv_test_core \
--gtest_filter='*Reduce*:*reduce*' \
--test_threads=1
```
### Performance Testing
Performance was measured with:
```bash
./bin/opencv_perf_core \
--gtest_filter='*reduceR*' \
--perf_threads=1
```
Test environment:
- SpacemiT K3
- VLEN = 256
- GCC 14.3.0
- Release build
- Single thread
- 10 samples per case
Results were compared against an unmodified `5.x` baseline.
`CV_32FC1 REDUCE_SUM2` median execution time (ms):
| Size | Baseline | Patched | Speedup |
| --- | ---: | ---: | ---: |
| 640x480 | 0.20 | 0.08 | 2.50x |
| 1280x720 | 0.64 | 0.29 | 2.21x |
| 1920x1080 | 1.33 | 1.05 | 1.27x |
Speedups are approximate and calculated from the rounded benchmark
output.
### Co-authors
- Yang Wang <yangwang@iscas.ac.cn>
- Yuansheng <yuansheng@isrc.iscas.ac.cn>
### Pull Request Readiness Checklist
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
FP8 model support in DNN - #29834
ONNX coverage after this PR: 76.6%
co-authored by: @SavyaSanchi-Sharma
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
core: fold a*alpha + b*beta + gamma into the fused OP_ADDW kernel - #29953
Part of #29443 — "fuse `a*alpha + b*beta + gamma`, where alpha, beta and gamma are scalars, into
some new `OP_ADDW`". Builds on #29937, which is required for correctness (see below).
### Problem
`OP_ADDW` already computes `a*alpha + b*beta + gamma` as a single kernel over two `v_fma`, and
`emitBinary()` already knows how to emit it — but the string front-end never recognized the
pattern, so `cv::texpr()` always took the written-out path: two multiplies, two adds, three temp
buffers and four passes over the data.
```
{0}*2.0 + {1}*3.0 + 1.0 at CV_32F
insns=4 temps=3 buffers=3 insns=1 temps=0 buffers=0
0: mul(1, 4) -> 5 ==> 0: addWeighted(1, 2) -> 11
1: mul(2, 7) -> 8 params=[2, 3, 1]
2: add(5, 8) -> 9
3: add(9, 11) -> 13
```
### Fix
A peephole in `emitBinary()`: `a*alpha + b*beta` folds into one `OP_ADDW`, and a trailing scalar
folds into that instruction's gamma rather than costing another pass.
`emitBinary()` may wrap a multiply in casts — an integer array times a fractional scalar computes
in the float domain and lands back in the array's own type — so the matcher accepts the optional
widening and narrowing casts around it. That is what makes the common 8-bit blend fuse; without
it `{0}*0.7 + {1}*0.3` at `CV_8U` stays at eight instructions and seven temps.
Shapes that are not an addWeighted keep their own meaning: an `a*b` term has no scalar factor, a
per-channel constant cannot ride the params block, and `CV_Bool` has no `OP_ADDW` form.
**Dependency on #29937.** This retires the instructions it folds, exactly as the
`abs(x - y) -> absdiff` peephole does. Without the `pinned` flag added in #29937 it would
reintroduce that bug for `u = {0}*2.0; v = {1}*3.0; u + v`, where the named terms are still live.
### Semantics on integer types
The fused kernel evaluates at its own work precision, so intermediate results no longer saturate
at each step. On integer inputs the result changes — it now agrees with `cv::addWeighted`, which
is what the expression means. This is the same trade the existing `abs(x - y)` peephole documents
in `emitUnary()`: the saturation artifacts of the literal expansion are never the desired result.
### Performance
1920x1080, best of 5 runs of 50 iterations, same build rebuilt both ways on 03ae9eac50:
| expression | depth | before | after |
| --- | --- | --- | --- |
| `a*0.7 + b*0.3` | 8U | 0.646 ms | 0.159 ms |
| `a*2 + b*3 + 1` | 8U | 0.261 ms | 0.150 ms |
| `a*2.5 + b*-1.5 + 7` | 32F | 0.360 ms | 0.214 ms |
| `a*2 + b*3 + 1` | 32F | 0.213 ms | 0.155 ms |
| `a*b + b*2` (control, not fused) | 32F | 0.393 ms | 0.358 ms |
The control row runs identical code in both builds and still moves by ~9%, so run-to-run noise on
this machine is around 10% — treat the 32F rows as indicative and the 8-bit rows as the real
result. Timings include `cv::texpr()` re-parsing and recompiling the expression on every call, so
the kernel-level gain is larger than the totals suggest.
### Tests
28 parameterised cases — 7 depths crossed with 4 weight sets, including a zero gamma and negative
weights — assert the fused result matches `cv::addWeighted`. Two further tests cover the variants
(no gamma, scalar written first, leading gamma) and the shapes the peephole must decline. 1801
tests in the arithmetic and TExpr suites pass.
Also adds the missing `OP_ADDW` case to `opName()`, which this change makes visible in every dump
of such a program.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Remove deprecated CommaInitializer API - #29948
This goes hand in hand with https://github.com/opencv/opencv_contrib/pull/4217
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
core: fix overflow in borderInterpolate BORDER_WRAP path - #29884
### Description
Fixes an integer overflow/underflow bug in `cv::borderInterpolate()` when using `BORDER_WRAP`. For extreme values of `p` (e.g. `INT_MIN` or `INT_MAX`), the previous implementation performed the modulo operation directly on `int`, which can invoke undefined behavior on overflow and produce a result outside the valid `[0, len)` range.
The fix widens the intermediate calculation to `int64_t` before taking the modulo, then adjusts for negative results and narrows back to `int` only once the value is confirmed to be in range.
### Changes
- `modules/core/src/copy.cpp`: use 64-bit intermediate arithmetic in the `BORDER_WRAP` branch of `borderInterpolate()` to avoid overflow.
- `modules/core/test/test_misc.cpp`: add `Core_BorderInterpolate.wrap_no_overflow_29232`, a regression test that exercises `BORDER_WRAP` with `INT_MIN` and `INT_MAX` and asserts the result stays within `[0, len)`.
Fixes#29232
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test for the patch (added in test_misc.cpp); no performance test needed as this is a bug fix with negligible performance impact.
- [x] The feature is well documented and sample code can be built with the project CMake
core: rvv: optimize 8U row reduce sum with vwaddu.wv - #29923
### Summary
Optimize the RVV implementation of `reduceRowSum_8u32s` by using the
native widening add instruction `vwaddu.wv`.
The generic scalable-vector path expands each `u8m1` input vector into
two `u16m1` halves and performs two separate load/add/store sequences
for the `u16` accumulation buffer.
On RVV, an `u8m1` vector and an `u16m2` vector have the same number of
elements. This allows the implementation to use an `u16m2` accumulator
directly and combine widening and addition with `vwaddu.wv`:
```text
u16m2 = u16m2 + u8m1
```
This removes the explicit widening and low/high `u16` accumulator split,
reducing the vector operations in the hot accumulation loop.
The existing scalar tail and 256-row `u16`-to-`u32` flush logic remain
unchanged.
The implementation is VLEN-agnostic.
### Functional Testing
The Reduce tests were run with:
```bash
./bin/opencv_test_core \
--gtest_filter='*Reduce*:*reduce*'
```
### Performance Testing
Performance was measured with:
```bash
OPENCV_FOR_THREADS_NUM=1 ./bin/opencv_perf_core \
--gtest_filter='*Reduce*:*reduce*'
```
Test environment:
- SpacemiT K3 / X100
- RVV 1.0
- VLEN = 256
- GCC 14.3.0
- Release build
- Single thread
Results were compared against an unmodified `5.x` baseline.
Median execution time (ms):
| Size / Type | Baseline | RVV `vwaddu.wv` | Speedup |
| --- | ---: | ---: | ---: |
| 640x480 8UC1 | 0.09 | 0.05 | 1.80x |
| 640x480 8UC4 | 0.36 | 0.16 | 2.25x |
| 1280x720 8UC1 | 0.27 | 0.12 | 2.25x |
| 1280x720 8UC4 | 1.08 | 0.56 | 1.93x |
| 1920x1080 8UC1 | 0.62 | 0.29 | 2.14x |
| 1920x1080 8UC4 | 2.43 | 1.11 | 2.19x |
Geometric mean speedup: **~2.09x**.
`REDUCE_MIN`, `REDUCE_MAX`, and `REDUCE_SUM2` performance remains
essentially unchanged, indicating that the speedup is localized to the
optimized 8-bit row SUM path.
### Co-authors
- Yang Wang <yangwang@iscas.ac.cn>
- Yuansheng <yuansheng@isrc.iscas.ac.cn>
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
FileStorage's XML/YAML/JSON parsers recurse once per nesting level
with no depth limit, allowing a small crafted file to exhaust the
stack and crash the process with an uncatchable SIGSEGV (CWE-674).
Add a shared CV_PERSISTENCE_MAX_DEPTH constant and thread a depth
counter through parseValue (XML/YAML) and parseSeq/parseMap (JSON),
raising a catchable cv::Exception via CV_PARSE_ERROR_CPP once the
limit is exceeded.
Fixes#29939
TExpr::moveToOutput() and the abs(x - y) -> absdiff(x, y) peephole in
TExpr::emitUnary() retire a value's arg slot (reclassify it to NONE) once
its single consumer has been emitted. That holds for an anonymous
intermediate, but a value the parser bound to a name ("t = ...;") may be
referenced again: the parser's name table still points at the retired slot,
so the later reference resolved to the reserved empty operand and
cv::texpr() silently returned a wrong result - for
t = {0} - {1}; abs(t) + t
t = {0} - {1}; (abs(t), t)
t = {0} + {1}; (t, t)
the reused name yielded input {0} instead of its own value, with no
assertion.
Mark a slot as pinned when the parser binds it to a name and skip both
retire manoeuvres for a pinned slot; each then takes the non-destructive
path it already has - moveToOutput() copies into the output via OP_CAST,
and the abs peephole falls through to the plain absdiff(a, 0) form, keeping
the OP_SUB that the name still needs. Anonymous intermediates are
unaffected, so the zero-temp fast path for single-op programs still fires.
core: fix RVV widening load lane count (#29907) - #29921
# core: fix RVV widening load lane count
Partially fixes#29907.
The RVV `v_load_expand` implementation used the source vector lane count for the widening load. For widening loads, the number of loaded elements must match the destination widened vector lane count instead.
This change:
* uses `VTraits<_Tpwvec>::vlanes()` for the RVV widening load and conversion;
* adds regression coverage for `texpr` `select()` with byte masks across multiple data types, channel counts, mask types, strided matrices, and in-place output.
### Pull Request Readiness Checklist
See details at the OpenCV contribution guidelines.
* [x] I agree to contribute to the project under Apache 2 License.
* [x] To the best of my knowledge, the proposed patch is not based on code under GPL or another license that is incompatible with OpenCV.
* [x] The PR is proposed to the proper branch.
* [x] There is a reference to the original bug report and related work.
* [x] There is an accuracy/regression test where applicable.
* [x] The change does not require documentation or sample updates.
core: handle zero-sized broadcast dimensions - #29911
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
Port the fix from #29878 to the 5.x branch.
This adds a guard for zero-sized destination matrices in
cv::broadcast() and a regression test covering broadcasting
from {1, 0} to {3, 0}.
The relevant BroadcastTo.* tests pass locally.
Related: #29878
core: fix addWeighted null kernel crash for f64 dtype and bool inputs - #29883Fixes#29880.
`cv::addWeighted` segfaults for `CV_8U`, `CV_8S`, `CV_16U`, `CV_16S`, `CV_16F`, `CV_16BF` and `CV_32F` inputs with `dtype=CV_64F`, and for `CV_Bool` inputs with any dtype. When no direct `T -> rdepth` kernel exists, `TExpr::emitBinary()` picks a wide work type and looks the kernel up again, but for those input types only `T -> T` and `T -> f32` kernels are generated, so the second lookup returns a null function pointer too. The `addInsn()` overload that takes an already resolved kernel stores it without checking, and `runInsn()` then calls through the null pointer.
Cast the operands to the work type when there is no kernel for them either, so the f64 (or f32) kernel runs on widened inputs. That is also what 4.x did, it converted the sources to the working type before computing, so an f64 destination keeps full precision instead of going through an f32 intermediate. Added the `CV_Assert` on the resolved kernel that the other emit paths already carry.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch (`5.x`, the element-wise engine this regressed in does not exist on `4.x`)
- [x] There is a reference to the original bug report and related work (#29880, regressed by #29426)
- [x] There is an accuracy test (`Core_Arithm.addWeighted_dtype_29880`, which segfaults without the fix); not applicable: performance test and opencv_extra test data
- [x] N/A: this is a bug fix, no new public API or documentation needed