Commit Graph
6091 Commits
Author SHA1 Message Date
Teddy-Yangjiale 95e8141250 Merge pull request #30069 from Teddy-Yangjiale:rvv-f64-arithm
core: enable CV_64F arithmetic SIMD kernels for scalable SIMD - #30069

### Summary

The CV_64F element-wise arithmetic kernels in `modules/core/src/arithm.simd.hpp` are guarded by `CV_SIMD_64F`, which is 0 on scalable-vector (RVV) builds, while the kernels themselves are already written with width-independent universal intrinsics (`v_float64`, `vx_load`, `VTraits<>::vlanes()`). On RVV every CV_64F arithmetic operation therefore falls back to the scalar kernels even though `CV_SIMD_SCALABLE_64F` is 1 and `v_float64` is fully supported.

The same file already uses the dual guard `CV_SIMD_64F || CV_SIMD_SCALABLE_64F` in four other places (see the comment at L166: "scalable (RVV) has v_float64 without CV_SIMD_64F"), so the 12 remaining single-guard sites look like an oversight. This PR widens them.

### What becomes vectorized on RVV (VLEN=256 => 4 double lanes)

- `cv::add` / `cv::subtract` (CV_64F)
- `cv::multiply` (CV_64F, plus the 32U/32S f64-work-type kernels)
- `cv::divide` (DIVW group: 32U/32S/64U/64S/64F)
- `cv::min` / `cv::max` (CV_64F)
- `cv::absdiff` (CV_64F)
- `cv::compare` (CV_64F)
- `cv::addWeighted` (AWD group: 32U/32S/64U/64S/64F)

All affected kernels are element-wise with no cross-lane reductions, so results are identical to the scalar kernels (per-element IEEE operations). The only theoretical difference is that the addWeighted vector body uses fused `v_fma` while the scalar tail does separate mul+add (at most 1 ulp); the same code already passes x86 AVX2/AVX512 CI where FMA is used as well.

On fixed-width backends (x86, AArch64, LoongArch, MIPS MSA, Power VSX) `CV_SIMD_64F` is already 1, so the guard change compiles to exactly the same code as before — no impact on other architectures. ARMv7/AArch32 and WASM SIMD128 have no f64 vector type at all (`CV_SIMD128_64F=0` by design) and keep the scalar path unchanged.

### Scope

Generic `rv64gc + CPU_DISPATCH=RVV` builds do not benefit yet (`arithm` is not a dispatched file); the win applies to builds where the core library itself is compiled with RVV enabled (`CPU_BASELINE=RVV`, the common embedded configuration on SpacemiT K1/K3 boards).

### Build configuration

Cross-compiled for SpacemiT K1 (8x 1.6 GHz, RVV 1.0, VLEN=256), GCC 14.2.1, `-O3`, both sides built with identical flags:

```bash
cmake -GNinja -DCMAKE_TOOLCHAIN_FILE=riscv64-spacemit-k1.toolchain.cmake \
      -DCMAKE_BUILD_TYPE=Release -DCPU_BASELINE=RVV -DCPU_DISPATCH= \
      -DBUILD_LIST=core,ts,calib,geometry -DBUILD_TESTS=ON -DBUILD_PERF_TESTS=ON \
      -DWITH_LAPACK=OFF -DWITH_EIGEN=OFF -DWITH_OPENCL=OFF \
      -DWITH_TIFF=OFF -DWITH_OPENJPEG=OFF -DWITH_JASPER=OFF -DWITH_OPENEXR=OFF \
      -DWITH_GDAL=OFF -DWITH_GDCM=OFF -DWITH_AVIF=OFF -DWITH_JPEGXL=OFF -DWITH_IMGCODEC_GIF=OFF ..

ninja opencv_core opencv_test_core opencv_perf_core opencv_test_calib opencv_test_geometry
```


### Correctness validation (on board)

```bash
export OPENCV_TEST_DATA_PATH=.../opencv_extra/testdata
export OPENCV_OPENCL_RUNTIME=disabled
./bin/opencv_test_core     --test_threads=4
./bin/opencv_test_calib    --test_threads=4
./bin/opencv_test_geometry --test_threads=4
```

| suite | baseline | with this PR |
|---|---|---|
| `opencv_test_core` (16757 tests) | 16756 pass, 1 fail: `Samples.findFile` (missing samples data path, environment) | **identical failure set** |
| `opencv_test_calib` (34 tests) | 34/34 pass | 34/34 pass |
| `opencv_test_geometry` (323 tests) | 323/323 pass | 323/323 pass |

Zero new failures; the arithm accuracy suites (`Core_*ElemWiseTest`, `Core_AddWeighted*`, `Core_*Mixed/ArithmMixedTest`, compare) exercise the newly vectorized kernels at all depths and pass unchanged.

### Performance (SpacemiT K1, single-threaded)

This PR adds a CV_64F perf fixture (`F64ArithmTest`: 9 ops x 4 sizes x CV_64FC1 = 36 cases) to `modules/core/perf/perf_arithm.cpp`, since the existing `BinaryOpTest` fixture does not parametrize CV_64F (and has no `compare`/`addWeighted` cases at all). Numbers below are per-case minima over 4 alternating A/B rounds x 50 samples, reported with the standard `modules/ts/misc/summary.py` (`-m min`).

With the default build flags (`-O3`, GCC 14.2.1):

```
Min (ms)

                 Name of Test                   base   cand     cand
                                                min4   min4     min4
                                                                 vs
                                                                base
                                                                min4
                                                             (x-factor)
absdiff::F64ArithmTest::(127x61, 64FC1)        0.035  0.026     1.33
absdiff::F64ArithmTest::(640x480, 64FC1)       1.580  1.644     0.96
absdiff::F64ArithmTest::(1280x720, 64FC1)      4.533  5.049     0.90
absdiff::F64ArithmTest::(1920x1080, 64FC1)     10.059 10.936    0.92
add::F64ArithmTest::(127x61, 64FC1)            0.027  0.028     0.98
add::F64ArithmTest::(640x480, 64FC1)           1.529  1.561     0.98
add::F64ArithmTest::(1280x720, 64FC1)          4.560  4.450     1.02
add::F64ArithmTest::(1920x1080, 64FC1)         9.544  9.549     1.00
addWeighted::F64ArithmTest::(127x61, 64FC1)    0.033  0.028     1.14
addWeighted::F64ArithmTest::(640x480, 64FC1)   1.582  1.581     1.00
addWeighted::F64ArithmTest::(1280x720, 64FC1)  4.579  4.797     0.95
addWeighted::F64ArithmTest::(1920x1080, 64FC1) 10.255 9.885     1.04
compare::F64ArithmTest::(127x61, 64FC1)        0.099  0.035     2.85
compare::F64ArithmTest::(640x480, 64FC1)       3.834  1.300     2.95
compare::F64ArithmTest::(1280x720, 64FC1)      11.456 3.847     2.98
compare::F64ArithmTest::(1920x1080, 64FC1)     25.670 8.738     2.94
divide::F64ArithmTest::(127x61, 64FC1)         0.064  0.065     0.98
divide::F64ArithmTest::(640x480, 64FC1)        2.493  2.352     1.06
divide::F64ArithmTest::(1280x720, 64FC1)       7.359  7.192     1.02
divide::F64ArithmTest::(1920x1080, 64FC1)      16.306 15.548    1.05
max::F64ArithmTest::(127x61, 64FC1)            0.029  0.026     1.12
max::F64ArithmTest::(640x480, 64FC1)           1.583  1.600     0.99
max::F64ArithmTest::(1280x720, 64FC1)          4.746  4.963     0.96
max::F64ArithmTest::(1920x1080, 64FC1)         9.827  10.143    0.97
min::F64ArithmTest::(127x61, 64FC1)            0.029  0.026     1.10
min::F64ArithmTest::(640x480, 64FC1)           1.625  1.630     1.00
min::F64ArithmTest::(1280x720, 64FC1)          4.821  5.000     0.96
min::F64ArithmTest::(1920x1080, 64FC1)         9.363  11.126    0.84
multiply::F64ArithmTest::(127x61, 64FC1)       0.029  0.027     1.10
multiply::F64ArithmTest::(640x480, 64FC1)      1.680  1.762     0.95
multiply::F64ArithmTest::(1280x720, 64FC1)     4.835  4.929     0.98
multiply::F64ArithmTest::(1920x1080, 64FC1)    10.427 10.564    0.99
subtract::F64ArithmTest::(127x61, 64FC1)       0.027  0.026     1.06
subtract::F64ArithmTest::(640x480, 64FC1)      1.605  1.660     0.97
subtract::F64ArithmTest::(1280x720, 64FC1)     4.758  5.016     0.95
subtract::F64ArithmTest::(1920x1080, 64FC1)    11.156 10.594    1.05
```

x-factor geomean **1.13x**. `compare` is a robust ~2.9x; most other ops land near parity because GCC already auto-vectorizes the simple scalar f64 loops at `-O3` on RVV (exactly the issue #30066 works around with `-fno-tree-vectorize`). Per-case values near parity swing with the known K1 bimodal timing behavior.

With `-fno-tree-vectorize` added to the build flags (the direction proposed in #30066 for RVV), the same A/B shows the kernel-level gains directly:

```
Min (ms)

                 Name of Test                    nt     nt       nt
                                                base   cand     cand
                                                min4   min4     min4
                                                                 vs
                                                                 nt
                                                                base
                                                                min4
                                                             (x-factor)
absdiff::F64ArithmTest::(127x61, 64FC1)        0.087  0.026     3.41
absdiff::F64ArithmTest::(640x480, 64FC1)       3.351  1.423     2.36
absdiff::F64ArithmTest::(1280x720, 64FC1)      10.189 4.595     2.22
absdiff::F64ArithmTest::(1920x1080, 64FC1)     22.476 10.094    2.23
add::F64ArithmTest::(127x61, 64FC1)            0.037  0.027     1.35
add::F64ArithmTest::(640x480, 64FC1)           1.538  1.446     1.06
add::F64ArithmTest::(1280x720, 64FC1)          4.644  4.403     1.05
add::F64ArithmTest::(1920x1080, 64FC1)         10.316 9.629     1.07
addWeighted::F64ArithmTest::(127x61, 64FC1)    0.076  0.027     2.81
addWeighted::F64ArithmTest::(640x480, 64FC1)   2.978  1.363     2.18
addWeighted::F64ArithmTest::(1280x720, 64FC1)  9.141  4.118     2.22
addWeighted::F64ArithmTest::(1920x1080, 64FC1) 19.856 9.075     2.19
compare::F64ArithmTest::(127x61, 64FC1)        0.103  0.031     3.31
compare::F64ArithmTest::(640x480, 64FC1)       3.993  1.225     3.26
compare::F64ArithmTest::(1280x720, 64FC1)      11.801 3.596     3.28
compare::F64ArithmTest::(1920x1080, 64FC1)     26.611 8.125     3.28
divide::F64ArithmTest::(127x61, 64FC1)         0.156  0.065     2.41
divide::F64ArithmTest::(640x480, 64FC1)        6.095  2.330     2.62
divide::F64ArithmTest::(1280x720, 64FC1)       18.191 6.880     2.64
divide::F64ArithmTest::(1920x1080, 64FC1)      40.743 15.331    2.66
max::F64ArithmTest::(127x61, 64FC1)            0.085  0.025     3.36
max::F64ArithmTest::(640x480, 64FC1)           3.272  1.428     2.29
max::F64ArithmTest::(1280x720, 64FC1)          10.009 4.581     2.18
max::F64ArithmTest::(1920x1080, 64FC1)         21.962 10.130    2.17
min::F64ArithmTest::(127x61, 64FC1)            0.086  0.025     3.40
min::F64ArithmTest::(640x480, 64FC1)           3.316  1.417     2.34
min::F64ArithmTest::(1280x720, 64FC1)          10.133 4.583     2.21
min::F64ArithmTest::(1920x1080, 64FC1)         22.334 10.160    2.20
multiply::F64ArithmTest::(127x61, 64FC1)       0.062  0.024     2.61
multiply::F64ArithmTest::(640x480, 64FC1)      2.430  1.362     1.78
multiply::F64ArithmTest::(1280x720, 64FC1)     7.427  4.344     1.71
multiply::F64ArithmTest::(1920x1080, 64FC1)    16.160 9.370     1.72
subtract::F64ArithmTest::(127x61, 64FC1)       0.038  0.025     1.48
subtract::F64ArithmTest::(640x480, 64FC1)      1.696  1.523     1.11
subtract::F64ArithmTest::(1280x720, 64FC1)     4.794  4.805     1.00
subtract::F64ArithmTest::(1920x1080, 64FC1)    12.058 10.820    1.11
```

x-factor geomean **2.09x** (36 cases, 1.00-3.41x): compare ~3.3x, divide ~2.6x, min/max ~2.2x, absdiff ~2.2x, addWeighted ~2.2x, multiply ~1.7-2.6x. `add`/`subtract` stay near parity in both regimes — they appear to be served by a different (already-vectorized) path before the kernel table is consulted.



### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
2026-09-29 08:31:04 +03:00
Ziyuan_Li 912e6863df Merge pull request #30075 from ziyuanLi-alex:rvv-transpose2d
core: RVV HAL transpose2d for 3-channel element sizes - #30075

### Summary

`cv::transpose` for 3-channel types (`8UC3`/`8SC3`, `16UC3`/`16SC3`/`16FC3`, `32SC3`/`32FC3`) falls back to the scalar core path on RISC-V. The RVV HAL `transpose2d` dispatches on *element size* and only implements `esz ∈ {1, 2, 4, 8}`. This PR adds kernels for
`esz = 3, 6, 12`, and an accuracy test that covers the new paths including non-continuous ROI.

### Root cause

- `hal/riscv-rvv/src/core/transpose.cpp`: the `Transpose2dFunc tab[]` table has entries only at indices 1/2/4/8, so `cv_hal_transpose2d` returns `CV_HAL_ERROR_NOT_IMPLEMENTED` for every 3-channel type and the caller runs the generic path.
- On the scalable RVV backend (`intrin_rvv_scalable.hpp`) `CV_SIMD128` is never defined, so the existing `transpose_8/16/32/48bit_simd` fast paths `modules/core/src/matrix_transform.cpp` are compiled out entirely; what runs is the scalar 4x4 element loop. 8UC3 (the default color type) is the most visible case.

### Implementation

The new kernels deinterleave the three channels of several source rows with `vlseg3e{L}` and emit the transposed rows with strided segment stores (`vssseg{6,8}e{L}`). The number of source rows per block was picked experimentally: 8 rows for 8-bit lanes, 2 rows for 16/32-bit lanes. For 16/32-bit lanes, taller blocks e.g. 8 rows will add register pressure and cause regression. A one-row tail covers the remainder; unaligned 16/32-bit inputs fall back to the generic path.

### Performance

SpacemiT K1 / X60, OpenCV 5.x, `--perf_force_samples=20
--perf_min_samples=20`, `BinaryOpTest.transpose2d`:

| type (esz) | 640x480 | 1280x720 | 1920x1080 |
|---|---:|---:|---:|
| `CV_8UC3` (3)  | 2.758 -> 1.648 ms (**1.67x**) | 16.772 -> 10.575 ms (**1.59x**) | 42.901 -> 39.552 ms (1.09x) |
| `CV_16SC3` (6) | 8.384 -> 3.556 ms (**2.36x**) | 31.112 -> 19.450 ms (**1.60x**) | 71.921 -> 62.460 ms (1.15x) |

Untouched element sizes (esz 1/2/4/8, 16 cases at 640x480 + 1280x720): geomean **1.005x**,
range 0.977x ... 1.058x.


### Accuracy

New test `Core_Transpose.C3ElementSizesWithRoi` (`modules/core/test/test_mat.cpp`) sweeps
`CV_8UC3, CV_8UC(6), CV_16SC3, CV_16FC3, CV_8UC(12), CV_32FC3` over sizes `1x1, 2x3, 137x5, 133x4` through non-continuous ROI views, byte-compares every element against the source, and asserts the bytes outside the destination ROI are untouched. On K1:

```
[  PASSED  ] 4 tests.   # Core_Transpose.C3ElementSizesWithRoi
                        # Core_Transpose/ElemWiseTest.accuracy/0
                        # Core_Rotate/ElemWiseTest.accuracy/0
```

`BinaryOpTest.transpose2d` already provides the performance coverage; no `opencv_extra` data is needed.


### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
2026-09-29 08:29:36 +03:00
Alexander Smorkalov 8e10da8b2d Fixed accuracy issue in reduce_sum with RISC-V RVV. 2026-09-28 16:39:17 +03:00
Alexander Smorkalov ef20307d84 Merge pull request #30071 from pranayr710:feat/hal-v-select-64bit
core: hal: add v_select for the 64-bit integer lanes
2026-09-28 09:49:03 +03:00
Abhishek Gola f34951afc4 Merge pull request #29844 from abhishek-gola:fix-avxvnni-assembler-check
Verify assembler can encode AVX-VNNI before enabling it - #29844

Closes: https://github.com/opencv/opencv/issues/29840

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
2026-09-28 08:35:16 +03:00
pranayr710 972c61904e core: hal: add v_select for the 64-bit integer lanes
v_select had no 64-bit integer form on any backend except WASM, so code
that needs to blend v_int64/v_uint64 through a comparison mask had to
apply the mask by hand. intrin_sse.hpp carried the two entries commented
out as TBD; the other backends simply stopped at 32-bit and float64.

Every backend already blends bitwise or through a byte/boolean cast, so
the lane width does not change the operation:

  SSE       _mm_blendv_pd under CV_SSE4_1, the xor/and/xor form otherwise
  AVX       _mm256_blendv_epi8, as the narrower types already use
  NEON      vbslq_u64 / vbslq_s64
  MSA       msa_bslq_u8 over the byte reinterpretation
  VSX       vec_sel with the same boolean cast v_float64x2 uses
  LSX/LASX  __lsx_vbitsel_v / __lasx_xvbitsel_v on the raw register
  RVV 0.7.1 vmerge_vvm with the b64 mask
  RVV       __riscv_vmerge, as for the narrower types

WASM already had both and is unchanged, and the generic v_reg
implementation in intrin_cpp.hpp already covers every type.

The 64-bit test chains could not simply call test_mask(), because its
v_signmask() expectations assume more than two lanes. Added test_select(),
which exercises v_select alone, and called it from TheTest<v_uint64>() and
TheTest<v_int64>().
2026-09-27 10:22:28 +05:30
Alexander Smorkalov e6376887da Merge pull request #29987 from Xlawy:opt/enable-scalable-vblas
core: enable VBLAS helpers for scalable SIMD
2026-09-25 10:00:15 +03:00
Alexander Smorkalov c57b34013c Merge pull request #30037 from Teddy-Yangjiale:rvv-gemm
core: enable order-preserving SIMD GEMM kernels for scalable SIMD
2026-09-25 08:53:58 +03:00
Vincent Rabaud 32b08d2b9e Remove references to C++ < 17 and older compiler versions.
According to https://github.com/opencv/opencv/wiki/OpenCV-4-to-5-migration#1-build-requirements
are not supported:
- GCC < 7
- clang < 9
- MSVC < 2017 (19.14)
2026-09-24 11:02:15 +02:00
Alexander Smorkalov 88061c6a75 Merge pull request #30052 from vrabaud:fallthrough
Enable -Wimplicit-fallthrough
2026-09-24 08:55:57 +03:00
Vincent Rabaud 0de5a7a266 Enable -Wimplicit-fallthrough 2026-09-23 14:28:55 +02:00
Vincent Rabaud ff3c9195f7 Remove now unused GET_OPTIMIZED 2026-09-23 11:18:41 +02:00
Vincent Rabaud 747ffc57be Merge pull request #30040 from vrabaud:function_ptr
Fix function pointer signature mismatches - #30040

Contrib PR: https://github.com/opencv/opencv_contrib/pull/4224

Calling a function through a function pointer with a mismatched signature is undefined behavior in C/C++ and causes Clang Control Flow Integrity to trap with `SIGILL` (`ud1`) at indirect call sites.

This is a follow-up on https://github.com/opencv/opencv/pull/28939

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
2026-09-23 10:44:20 +03:00
Alexander Smorkalov 924f87f481 Merge pull request #30018 from Xlawy:opt/rvv-integer-norm-mask-reuse
RVV: reuse mask chunks across integer norm channels
2026-09-23 09:31:13 +03:00
Alexander Smorkalov 112ca450c5 Merge pull request #30034 from asmorkalov:as/flip_alignment
Take memory aligment into account in flip.
2026-09-22 13:12:36 +03:00
Teddy-Yangjiale c0a18471a0 core: enable order-preserving SIMD GEMM kernels for scalable SIMD 2026-09-22 16:11:59 +08:00
Alexander Smorkalov 57a557ce30 Take memory aligment into account in flip. 2026-09-22 10:00:12 +03:00
Alexander Smorkalov 2e72f5b1de Update TExpr inplace check to handle float arithmetics optimizations. 2026-09-22 09:31:13 +03:00
Alexander Smorkalov f2f470a455 Merge pull request #30008 from pbkx:fix-calc-covar-vector-mean-roi-5x
fix calcCovarMatrix vector mean ROI handling
2026-09-22 08:11:18 +03:00
Alexander Smorkalov 8e719aaeff Merge pull request #30028 from asmorkalov:as/texpr_mem_alignment2
Fixed more memory alignment issues in TEexpr
2026-09-22 08:09:30 +03:00
Alexander Smorkalov b6718c0ff7 Merge pull request #30032 from lrycro:fix/glob-readdir-leak-5x
core: fix memory leak in glob()'s readdir() on WinRT/_WIN32_WCE
2026-09-22 08:09:04 +03:00
Alexander Smorkalov c9ef3617a9 Merge pull request #30020 from vrabaud:eigen
Add Eigen conversions for Affine3 and Quat
2026-09-21 19:22:54 +03:00
Sewon Ahn 676c0c9e04 core: fix memory leak in glob()'s readdir() on WinRT/_WIN32_WCE
readdir() allocates a new buffer for dir->ent.d_name on every call
and overwrites the previous pointer without freeing it. Since
cv::glob() calls readdir() once per directory entry, every call
except the last leaks its allocation. Under _WIN32_WCE, DIR has no
destructor at all, so every allocation leaks, including the last
one.

(cherry picked from commit 20b9529bc3)
2026-09-22 01:01:01 +09:00
Alexander Smorkalov 804901f13d Fixed more memory alignment issues. 2026-09-21 16:22:40 +03:00
Vincent Rabaud 3a2e5150a3 Add Eigen conversions for Affine3 and Quat 2026-09-21 11:14:12 +02:00
6c29da7a81 rvv: reuse mask chunks across integer norm channels
Process masked integer norms chunk by chunk and reuse each loaded mask
and predicate across channels, avoiding repeated mask processing.

Extend norm_mask performance coverage with three-channel 8U, 8S, 16U,
16S, and 32S inputs for INF, L1, and L2 norms.

Co-authored-by: Yang Wang <yangwang@iscas.ac.cn>
Co-authored-by: Yuansheng <yuansheng@isrc.iscas.ac.cn>
2026-09-21 14:18:08 +08:00
pbkx dadaa2dc25 fix InputArray empty for vector<UMat> 2026-09-20 17:32:53 -07:00
pbkx c5b17b6ccd fix calcCovarMatrix vector mean ROI handling 2026-09-20 15:54:54 -07:00
Muditya Raghav 3b580381f1 Merge pull request #29981 from 0xMudit:doc-mat-type-bit-layout
doc: document the bit layout of Mat::type() - #29981

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
- [ ] The feature is well documented and sample code can be built with the project CMake

### Description

Fixes #24901.

#### Problem

`cv::Mat::type()` returns a packed bit-field, but the encoding was not documented. Users
had to reverse-engineer the layout from the `CV_*` macros to answer questions such as how many
channels fit, how to build a type from a depth and channel count, and which bits are reserved for
the matrix flags.

#### Change

Expanded the Doxygen for `Mat::type()` in `modules/core/include/opencv2/core/mat.hpp` to describe
the layout used by the 5.x branch:

- bits 0-4 (`CV_MAT_DEPTH_MASK`) – element depth (5 bits);
- bits 5-11 (`CV_MAT_CN_MASK`) – number of channels minus one (7 bits), i.e. 1..`CV_CN_MAX` (128);
- together these occupy the lowest 12 bits (`CV_MAT_TYPE_MASK`).

The description also points to `CV_MAT_DEPTH()`, `CV_MAT_CN()` and `CV_MAKETYPE()`, and notes that
the continuity (`CV_MAT_CONT_FLAG`) and submatrix (`CV_SUBMAT_FLAG`) bits of `Mat::flags` are not
part of the returned value.

#### Branch note

This PR targets `5.x` only. The encoding changed between branches: 5.x uses `CV_CN_SHIFT == 5`
(5-bit depth, 7-bit channel count), whereas 4.x uses `CV_CN_SHIFT == 3`. As requested in the issue,
a separate `4.x` PR would be needed for that branch.

#### Verification

Documentation-only change; no code or behavior is modified. The bit ranges and macro names were
checked against `modules/core/include/opencv2/core/hal/interface.h` and
`modules/core/include/opencv2/core/cvdef.h`:

```
CV_CN_MAX            128
CV_CN_SHIFT          5
CV_DEPTH_MAX         (1 << CV_CN_SHIFT)          // 32
CV_MAT_DEPTH_MASK    (CV_DEPTH_MAX - 1)          // 0x1F   -> bits 0-4
CV_MAT_CN_MASK       ((CV_CN_MAX - 1) << 5)      // 0xFE0  -> bits 5-11
CV_MAT_TYPE_MASK     (CV_DEPTH_MAX*CV_CN_MAX-1)  // 0xFFF
```
2026-09-19 14:30:13 +03:00
ebf4eb3dad core: enable VBLAS helpers for scalable SIMD
Extend the CV_SIMD guard in lapack.cpp to include CV_SIMD_SCALABLE,
addressing the TODO updated in PR24325.

Co-authored-by: Yang Wang <yangwang@iscas.ac.cn>
Co-authored-by: Yuansheng <yuansheng@isrc.iscas.ac.cn>
2026-09-18 20:23:45 +08:00
zhangjinhan fb96a94a03 Merge pull request #29930 from Xlawy:opt/rvv-reduce-sum2-32f
core: optimize float REDUCE_SUM2 with RVV - #29930

### Summary

Add an RVV 1.0 kernel for `cv::reduce` with `dim=0` and `CV_32F`
input/output through the CPU dispatch mechanism.

The optimized path targets the floating-point `REDUCE_SUM2` operation.

### Implementation

Process four source rows per vector iteration to reduce intermediate
buffer traffic while preserving source-row accumulation order.

The RVV kernel uses LMUL=4 and dynamic vector lengths for tail handling.

Output writes are deferred until all input rows have been read. This
preserves correct behavior when the source and destination matrices
overlap.

The implementation is VLEN-agnostic.

### Functional Testing

The Reduce tests were run with:

```bash
./bin/opencv_test_core \
    --gtest_filter='*Reduce*:*reduce*' \
    --test_threads=1
```

### Performance Testing

Performance was measured with:

```bash
./bin/opencv_perf_core \
    --gtest_filter='*reduceR*' \
    --perf_threads=1
```

Test environment:

- SpacemiT K3
- VLEN = 256
- GCC 14.3.0
- Release build
- Single thread
- 10 samples per case

Results were compared against an unmodified `5.x` baseline.

`CV_32FC1 REDUCE_SUM2` median execution time (ms):

| Size | Baseline | Patched | Speedup |
| --- | ---: | ---: | ---: |
| 640x480 | 0.20 | 0.08 | 2.50x |
| 1280x720 | 0.64 | 0.29 | 2.21x |
| 1920x1080 | 1.33 | 1.05 | 1.27x |

Speedups are approximate and calculated from the rounded benchmark
output.

### Co-authors

- Yang Wang <yangwang@iscas.ac.cn>
- Yuansheng <yuansheng@isrc.iscas.ac.cn>

### Pull Request Readiness Checklist

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
2026-09-18 14:56:16 +03:00
Alexander Smorkalov 84bb9b20c3 Merge pull request #29974 from asmorkalov:as/win_warning_fix
Warnings fix on Windows.
2026-09-17 19:46:09 +03:00
Alexander Smorkalov df75b5f967 Warnings fix on Windows. 2026-09-17 18:43:32 +03:00
Abhishek Gola b2b4f34820 Merge pull request #29834 from abhishek-gola:dnn-fp8-support
FP8 model support in DNN - #29834

ONNX coverage after this PR: 76.6%

co-authored by: @SavyaSanchi-Sharma 

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
2026-09-17 16:15:30 +03:00
pranayr710 d28b360004 Merge pull request #29953 from pranayr710:feat/texpr-fuse-addweighted
core: fold a*alpha + b*beta + gamma into the fused OP_ADDW kernel - #29953

Part of #29443 — "fuse `a*alpha + b*beta + gamma`, where alpha, beta and gamma are scalars, into
some new `OP_ADDW`". Builds on #29937, which is required for correctness (see below).

### Problem

`OP_ADDW` already computes `a*alpha + b*beta + gamma` as a single kernel over two `v_fma`, and
`emitBinary()` already knows how to emit it — but the string front-end never recognized the
pattern, so `cv::texpr()` always took the written-out path: two multiplies, two adds, three temp
buffers and four passes over the data.

```
{0}*2.0 + {1}*3.0 + 1.0   at CV_32F

  insns=4  temps=3  buffers=3          insns=1  temps=0  buffers=0
    0: mul(1, 4)  -> 5          ==>      0: addWeighted(1, 2) -> 11
    1: mul(2, 7)  -> 8                      params=[2, 3, 1]
    2: add(5, 8)  -> 9
    3: add(9, 11) -> 13
```

### Fix

A peephole in `emitBinary()`: `a*alpha + b*beta` folds into one `OP_ADDW`, and a trailing scalar
folds into that instruction's gamma rather than costing another pass.

`emitBinary()` may wrap a multiply in casts — an integer array times a fractional scalar computes
in the float domain and lands back in the array's own type — so the matcher accepts the optional
widening and narrowing casts around it. That is what makes the common 8-bit blend fuse; without
it `{0}*0.7 + {1}*0.3` at `CV_8U` stays at eight instructions and seven temps.

Shapes that are not an addWeighted keep their own meaning: an `a*b` term has no scalar factor, a
per-channel constant cannot ride the params block, and `CV_Bool` has no `OP_ADDW` form.

**Dependency on #29937.** This retires the instructions it folds, exactly as the
`abs(x - y) -> absdiff` peephole does. Without the `pinned` flag added in #29937 it would
reintroduce that bug for `u = {0}*2.0; v = {1}*3.0; u + v`, where the named terms are still live.

### Semantics on integer types

The fused kernel evaluates at its own work precision, so intermediate results no longer saturate
at each step. On integer inputs the result changes — it now agrees with `cv::addWeighted`, which
is what the expression means. This is the same trade the existing `abs(x - y)` peephole documents
in `emitUnary()`: the saturation artifacts of the literal expansion are never the desired result.

### Performance

1920x1080, best of 5 runs of 50 iterations, same build rebuilt both ways on 03ae9eac50:

| expression | depth | before | after |
| --- | --- | --- | --- |
| `a*0.7 + b*0.3` | 8U | 0.646 ms | 0.159 ms |
| `a*2 + b*3 + 1` | 8U | 0.261 ms | 0.150 ms |
| `a*2.5 + b*-1.5 + 7` | 32F | 0.360 ms | 0.214 ms |
| `a*2 + b*3 + 1` | 32F | 0.213 ms | 0.155 ms |
| `a*b + b*2` (control, not fused) | 32F | 0.393 ms | 0.358 ms |

The control row runs identical code in both builds and still moves by ~9%, so run-to-run noise on
this machine is around 10% — treat the 32F rows as indicative and the 8-bit rows as the real
result. Timings include `cv::texpr()` re-parsing and recompiling the expression on every call, so
the kernel-level gain is larger than the totals suggest.

### Tests

28 parameterised cases — 7 depths crossed with 4 weight sets, including a zero gamma and negative
weights — assert the fused result matches `cv::addWeighted`. Two further tests cover the variants
(no gamma, scalar written first, leading gamma) and the shapes the peephole must decline. 1801
tests in the arithmetic and TExpr suites pass.

Also adds the missing `OP_ADDW` case to `opName()`, which this change makes visible in every dump
of such a program.

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
2026-09-16 16:10:41 +03:00
Vincent Rabaud 551dbcac54 Merge pull request #29948 from vrabaud:comma_initializer
Remove deprecated CommaInitializer API - #29948

This goes hand in hand with https://github.com/opencv/opencv_contrib/pull/4217

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
2026-09-16 13:59:50 +03:00
Akshar Singhal 60e518784c Merge pull request #29884 from Aks27-hub:fix-borderwrap-overflow
core: fix overflow in borderInterpolate BORDER_WRAP path - #29884

### Description

Fixes an integer overflow/underflow bug in `cv::borderInterpolate()` when using `BORDER_WRAP`. For extreme values of `p` (e.g. `INT_MIN` or `INT_MAX`), the previous implementation performed the modulo operation directly on `int`, which can invoke undefined behavior on overflow and produce a result outside the valid `[0, len)` range.

The fix widens the intermediate calculation to `int64_t` before taking the modulo, then adjusts for negative results and narrows back to `int` only once the value is confirmed to be in range.

### Changes

- `modules/core/src/copy.cpp`: use 64-bit intermediate arithmetic in the `BORDER_WRAP` branch of `borderInterpolate()` to avoid overflow.
- `modules/core/test/test_misc.cpp`: add `Core_BorderInterpolate.wrap_no_overflow_29232`, a regression test that exercises `BORDER_WRAP` with `INT_MIN` and `INT_MAX` and asserts the result stays within `[0, len)`.

Fixes #29232

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test for the patch (added in test_misc.cpp); no performance test needed as this is a bug fix with negligible performance impact.
- [x] The feature is well documented and sample code can be built with the project CMake
2026-09-16 09:22:35 +03:00
Alexander Smorkalov 814692f9c8 Merge pull request #29957 from asmorkalov:as/getElemSize_doc
Added documentation for getElemSize
2026-09-15 15:19:29 +03:00
zhangjinhan 3a5032e190 Merge pull request #29923 from Xlawy:opt/rvv-reduce-8u-row-sum
core: rvv: optimize 8U row reduce sum with vwaddu.wv - #29923

### Summary

Optimize the RVV implementation of `reduceRowSum_8u32s` by using the
native widening add instruction `vwaddu.wv`.

The generic scalable-vector path expands each `u8m1` input vector into
two `u16m1` halves and performs two separate load/add/store sequences
for the `u16` accumulation buffer.

On RVV, an `u8m1` vector and an `u16m2` vector have the same number of
elements. This allows the implementation to use an `u16m2` accumulator
directly and combine widening and addition with `vwaddu.wv`:

```text
u16m2 = u16m2 + u8m1
```

This removes the explicit widening and low/high `u16` accumulator split,
reducing the vector operations in the hot accumulation loop.

The existing scalar tail and 256-row `u16`-to-`u32` flush logic remain
unchanged.

The implementation is VLEN-agnostic.

### Functional Testing

The Reduce tests were run with:

```bash
./bin/opencv_test_core \
    --gtest_filter='*Reduce*:*reduce*'
```

### Performance Testing

Performance was measured with:

```bash
OPENCV_FOR_THREADS_NUM=1 ./bin/opencv_perf_core \
    --gtest_filter='*Reduce*:*reduce*'
```

Test environment:

- SpacemiT K3 / X100
- RVV 1.0
- VLEN = 256
- GCC 14.3.0
- Release build
- Single thread

Results were compared against an unmodified `5.x` baseline.

Median execution time (ms):

| Size / Type | Baseline | RVV `vwaddu.wv` | Speedup |
| --- | ---: | ---: | ---: |
| 640x480 8UC1 | 0.09 | 0.05 | 1.80x |
| 640x480 8UC4 | 0.36 | 0.16 | 2.25x |
| 1280x720 8UC1 | 0.27 | 0.12 | 2.25x |
| 1280x720 8UC4 | 1.08 | 0.56 | 1.93x |
| 1920x1080 8UC1 | 0.62 | 0.29 | 2.14x |
| 1920x1080 8UC4 | 2.43 | 1.11 | 2.19x |

Geometric mean speedup: **~2.09x**.

`REDUCE_MIN`, `REDUCE_MAX`, and `REDUCE_SUM2` performance remains
essentially unchanged, indicating that the speedup is localized to the
optimized 8-bit row SUM path.

### Co-authors

- Yang Wang <yangwang@iscas.ac.cn>
- Yuansheng <yuansheng@isrc.iscas.ac.cn>

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
2026-09-15 15:18:17 +03:00
Alexander Smorkalov 6833d08418 Added documentation for getElemSize 2026-09-15 14:23:28 +03:00
Alexander Smorkalov 393917a998 Merge pull request #29945 from Rishiii57:fix/filestorage-recursion-depth-limit-29939
core(persistence): add recursion depth limit to XML/YAML/JSON parsers
2026-09-15 12:02:50 +03:00
Alexander Smorkalov 955634486a Merge pull request #29925 from asmorkalov:as/texpr_alignment_fix
Memory alignment fix in TExpr impl for RISC-V RVV
2026-09-14 12:26:31 +03:00
Rishiii57 56334789a3 core(persistence): add recursion depth limit to XML/YAML/JSON parsers
FileStorage's XML/YAML/JSON parsers recurse once per nesting level
with no depth limit, allowing a small crafted file to exhaust the
stack and crash the process with an uncatchable SIGSEGV (CWE-674).

Add a shared CV_PERSISTENCE_MAX_DEPTH constant and thread a depth
counter through parseValue (XML/YAML) and parseSeq/parseMap (JSON),
raising a catchable cv::Exception via CV_PARSE_ERROR_CPP once the
limit is exceeded.

Fixes #29939
2026-09-14 07:48:54 +05:30
pranayr710 78ed76f0d9 core: keep named texpr values alive across slot-retiring optimizations
TExpr::moveToOutput() and the abs(x - y) -> absdiff(x, y) peephole in
TExpr::emitUnary() retire a value's arg slot (reclassify it to NONE) once
its single consumer has been emitted. That holds for an anonymous
intermediate, but a value the parser bound to a name ("t = ...;") may be
referenced again: the parser's name table still points at the retired slot,
so the later reference resolved to the reserved empty operand and
cv::texpr() silently returned a wrong result - for

    t = {0} - {1}; abs(t) + t
    t = {0} - {1}; (abs(t), t)
    t = {0} + {1}; (t, t)

the reused name yielded input {0} instead of its own value, with no
assertion.

Mark a slot as pinned when the parser binds it to a name and skip both
retire manoeuvres for a pinned slot; each then takes the non-destructive
path it already has - moveToOutput() copies into the output via OP_CAST,
and the abs peephole falls through to the plain absdiff(a, 0) form, keeping
the OP_SUB that the name still needs. Anonymous intermediates are
unaffected, so the zero-temp fast path for single-op programs still fires.
2026-09-13 02:30:55 +05:30
Alexander Smorkalov 254269f094 Merge pull request #29877 from cuishuang:core-reject-invalid-bool
core: reject invalid boolean values in CommandLineParser
2026-09-11 15:43:44 +03:00
Alexander Smorkalov a87a82f7d8 Memory alignment fix in TExpr impl for RISC-V RVV 2026-09-10 16:06:32 +03:00
MUHAMMAD AHMAD MASOOD 1b973eb0d2 Merge pull request #29921 from ahmadmasood43:fix-29907-rvv-texpr-mask
core: fix RVV widening load lane count (#29907) - #29921

# core: fix RVV widening load lane count

Partially fixes #29907.

The RVV `v_load_expand` implementation used the source vector lane count for the widening load. For widening loads, the number of loaded elements must match the destination widened vector lane count instead.

This change:

* uses `VTraits<_Tpwvec>::vlanes()` for the RVV widening load and conversion;
* adds regression coverage for `texpr` `select()` with byte masks across multiple data types, channel counts, mask types, strided matrices, and in-place output.

### Pull Request Readiness Checklist

See details at the OpenCV contribution guidelines.

* [x] I agree to contribute to the project under Apache 2 License.
* [x] To the best of my knowledge, the proposed patch is not based on code under GPL or another license that is incompatible with OpenCV.
* [x] The PR is proposed to the proper branch.
* [x] There is a reference to the original bug report and related work.
* [x] There is an accuracy/regression test where applicable.
* [x] The change does not require documentation or sample updates.
2026-09-10 13:56:25 +03:00
Sridhar 9940db5599 Merge pull request #29911 from sridhar-git05:fix-broadcast-zero-dimension-5x
core: handle zero-sized broadcast dimensions - #29911

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake

Port the fix from #29878 to the 5.x branch.

This adds a guard for zero-sized destination matrices in
cv::broadcast() and a regression test covering broadcasting
from {1, 0} to {3, 0}.

The relevant BroadcastTo.* tests pass locally.

Related: #29878
2026-09-10 10:31:04 +03:00
cuishuang fd00f2b2d0 core: reject invalid boolean values in CommandLineParser 2026-09-09 20:12:34 +08:00
Lazizbek Ergashev 24c45b01ab Merge pull request #29883 from lazerg:fix/issue-29880-addweighted-null-kernel
core: fix addWeighted null kernel crash for f64 dtype and bool inputs - #29883

Fixes #29880.

`cv::addWeighted` segfaults for `CV_8U`, `CV_8S`, `CV_16U`, `CV_16S`, `CV_16F`, `CV_16BF` and `CV_32F` inputs with `dtype=CV_64F`, and for `CV_Bool` inputs with any dtype. When no direct `T -> rdepth` kernel exists, `TExpr::emitBinary()` picks a wide work type and looks the kernel up again, but for those input types only `T -> T` and `T -> f32` kernels are generated, so the second lookup returns a null function pointer too. The `addInsn()` overload that takes an already resolved kernel stores it without checking, and `runInsn()` then calls through the null pointer.

Cast the operands to the work type when there is no kernel for them either, so the f64 (or f32) kernel runs on widened inputs. That is also what 4.x did, it converted the sources to the working type before computing, so an f64 destination keeps full precision instead of going through an f32 intermediate. Added the `CV_Assert` on the resolved kernel that the other emit paths already carry.

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch (`5.x`, the element-wise engine this regressed in does not exist on `4.x`)
- [x] There is a reference to the original bug report and related work (#29880, regressed by #29426)
- [x] There is an accuracy test (`Core_Arithm.addWeighted_dtype_29880`, which segfaults without the fix); not applicable: performance test and opencv_extra test data
- [x] N/A: this is a bug fix, no new public API or documentation needed
2026-09-08 20:03:15 +03:00