core: add ARMPL HAL backend for cv::invert(DECOMP_LU) - #30080
## Summary
- Follow-up to #[30035](https://github.com/opencv/opencv/pull/30035) to apply the same `cv::invert` optimization to the 5.x branch.
- Adds an ARM Performance Library (ArmPL) HAL backend for cv::invert under DECOMP_LU (cv_hal_LU32f / cv_hal_LU64f), replacing OpenCV's default LU decomposition path with ArmPL's LAPACKE_sgesv / LAPACKE_dgesv for float and double matrices on AArch64.
## Changes
**hal/armpl/include/armpl_hal_core.hpp:**
- Declare armpl_hal_LU32f / armpl_hal_LU64f and register them as cv_hal_LU32f / cv_hal_LU64f.
**hal/armpl/src/armpl_hal_core.cpp:**
- Add armpl_lu() function (works for both float and double), that does the LU decomposition of the matrix:
- If a right-hand side b is given, then LAPACKE_sgesv/dgesv is called.
- If no right-hand side is given, it only factorizes the matrix using LAPACKE_sgetrf/dgetrf
- Fall back to OpenCV's default (non-HAL) implementation for small matrices (m < 100).
dnn: fold Shape of static-shape inputs in constFold() - #30065
When a graph input has a fully known shape, `constFold()` now replaces `Shape` ops on it with constants. The usual const-folding cascade then also removes the Gather/Concat/Cast/Slice ops that follow.
**What it adds**
- `graph_const_fold.cpp`: tracks known shapes and types through the graph during folding and folds `Shape` layers whose input shape is known.
- `net_impl2.cpp`: once a `Shape` has been folded, `setGraphInput()` rejects an input of a different shape instead of running with stale constants. It also fixes `tryInferGraphShapes()` to report untyped const or empty inputs as `Mat().type()`. The extra folding exposed a Resize type-check failure without this fix.
- `test_fusion.cpp`: two tests, one checking that no `Shape` layer remains after folding, one checking that a different input shape is rejected.
**Effect on graphs** (outputs identical to 5.x)
CRNN 84 → 34 layers, ViT/DeiT 361 → 274, BEiT 373 → 286, MobileViT_XS 352 → 289.
**Performance** (OCV/CPU, 32 threads; median of 5 interleaved runs per side)
| Model | 5.x, ms | PR, ms | Speedup |
|---|---|---|---|
| MobileViT_XS | 6.50 | 6.23 | 1.04x |
| BEiT_Base_Patch16_224 | 30.03 | 29.23 | 1.03x |
| VIT_Base_Patch16_224 | 27.17 | 26.53 | 1.02x |
| DeiT_Tiny_Patch16_224 | 4.90 | 4.81 | 1.02x |
| CRNN | 2.52 | 2.51 | 1.01x |
| FacePaint | 328.1 | 329.5 | 1.00x |
**Dependency**
This depends on the upcoming CUDA integration in DNN and should be merged after it.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
core: enable CV_64F arithmetic SIMD kernels for scalable SIMD - #30069
### Summary
The CV_64F element-wise arithmetic kernels in `modules/core/src/arithm.simd.hpp` are guarded by `CV_SIMD_64F`, which is 0 on scalable-vector (RVV) builds, while the kernels themselves are already written with width-independent universal intrinsics (`v_float64`, `vx_load`, `VTraits<>::vlanes()`). On RVV every CV_64F arithmetic operation therefore falls back to the scalar kernels even though `CV_SIMD_SCALABLE_64F` is 1 and `v_float64` is fully supported.
The same file already uses the dual guard `CV_SIMD_64F || CV_SIMD_SCALABLE_64F` in four other places (see the comment at L166: "scalable (RVV) has v_float64 without CV_SIMD_64F"), so the 12 remaining single-guard sites look like an oversight. This PR widens them.
### What becomes vectorized on RVV (VLEN=256 => 4 double lanes)
- `cv::add` / `cv::subtract` (CV_64F)
- `cv::multiply` (CV_64F, plus the 32U/32S f64-work-type kernels)
- `cv::divide` (DIVW group: 32U/32S/64U/64S/64F)
- `cv::min` / `cv::max` (CV_64F)
- `cv::absdiff` (CV_64F)
- `cv::compare` (CV_64F)
- `cv::addWeighted` (AWD group: 32U/32S/64U/64S/64F)
All affected kernels are element-wise with no cross-lane reductions, so results are identical to the scalar kernels (per-element IEEE operations). The only theoretical difference is that the addWeighted vector body uses fused `v_fma` while the scalar tail does separate mul+add (at most 1 ulp); the same code already passes x86 AVX2/AVX512 CI where FMA is used as well.
On fixed-width backends (x86, AArch64, LoongArch, MIPS MSA, Power VSX) `CV_SIMD_64F` is already 1, so the guard change compiles to exactly the same code as before — no impact on other architectures. ARMv7/AArch32 and WASM SIMD128 have no f64 vector type at all (`CV_SIMD128_64F=0` by design) and keep the scalar path unchanged.
### Scope
Generic `rv64gc + CPU_DISPATCH=RVV` builds do not benefit yet (`arithm` is not a dispatched file); the win applies to builds where the core library itself is compiled with RVV enabled (`CPU_BASELINE=RVV`, the common embedded configuration on SpacemiT K1/K3 boards).
### Build configuration
Cross-compiled for SpacemiT K1 (8x 1.6 GHz, RVV 1.0, VLEN=256), GCC 14.2.1, `-O3`, both sides built with identical flags:
```bash
cmake -GNinja -DCMAKE_TOOLCHAIN_FILE=riscv64-spacemit-k1.toolchain.cmake \
-DCMAKE_BUILD_TYPE=Release -DCPU_BASELINE=RVV -DCPU_DISPATCH= \
-DBUILD_LIST=core,ts,calib,geometry -DBUILD_TESTS=ON -DBUILD_PERF_TESTS=ON \
-DWITH_LAPACK=OFF -DWITH_EIGEN=OFF -DWITH_OPENCL=OFF \
-DWITH_TIFF=OFF -DWITH_OPENJPEG=OFF -DWITH_JASPER=OFF -DWITH_OPENEXR=OFF \
-DWITH_GDAL=OFF -DWITH_GDCM=OFF -DWITH_AVIF=OFF -DWITH_JPEGXL=OFF -DWITH_IMGCODEC_GIF=OFF ..
ninja opencv_core opencv_test_core opencv_perf_core opencv_test_calib opencv_test_geometry
```
### Correctness validation (on board)
```bash
export OPENCV_TEST_DATA_PATH=.../opencv_extra/testdata
export OPENCV_OPENCL_RUNTIME=disabled
./bin/opencv_test_core --test_threads=4
./bin/opencv_test_calib --test_threads=4
./bin/opencv_test_geometry --test_threads=4
```
| suite | baseline | with this PR |
|---|---|---|
| `opencv_test_core` (16757 tests) | 16756 pass, 1 fail: `Samples.findFile` (missing samples data path, environment) | **identical failure set** |
| `opencv_test_calib` (34 tests) | 34/34 pass | 34/34 pass |
| `opencv_test_geometry` (323 tests) | 323/323 pass | 323/323 pass |
Zero new failures; the arithm accuracy suites (`Core_*ElemWiseTest`, `Core_AddWeighted*`, `Core_*Mixed/ArithmMixedTest`, compare) exercise the newly vectorized kernels at all depths and pass unchanged.
### Performance (SpacemiT K1, single-threaded)
This PR adds a CV_64F perf fixture (`F64ArithmTest`: 9 ops x 4 sizes x CV_64FC1 = 36 cases) to `modules/core/perf/perf_arithm.cpp`, since the existing `BinaryOpTest` fixture does not parametrize CV_64F (and has no `compare`/`addWeighted` cases at all). Numbers below are per-case minima over 4 alternating A/B rounds x 50 samples, reported with the standard `modules/ts/misc/summary.py` (`-m min`).
With the default build flags (`-O3`, GCC 14.2.1):
```
Min (ms)
Name of Test base cand cand
min4 min4 min4
vs
base
min4
(x-factor)
absdiff::F64ArithmTest::(127x61, 64FC1) 0.035 0.026 1.33
absdiff::F64ArithmTest::(640x480, 64FC1) 1.580 1.644 0.96
absdiff::F64ArithmTest::(1280x720, 64FC1) 4.533 5.049 0.90
absdiff::F64ArithmTest::(1920x1080, 64FC1) 10.059 10.936 0.92
add::F64ArithmTest::(127x61, 64FC1) 0.027 0.028 0.98
add::F64ArithmTest::(640x480, 64FC1) 1.529 1.561 0.98
add::F64ArithmTest::(1280x720, 64FC1) 4.560 4.450 1.02
add::F64ArithmTest::(1920x1080, 64FC1) 9.544 9.549 1.00
addWeighted::F64ArithmTest::(127x61, 64FC1) 0.033 0.028 1.14
addWeighted::F64ArithmTest::(640x480, 64FC1) 1.582 1.581 1.00
addWeighted::F64ArithmTest::(1280x720, 64FC1) 4.579 4.797 0.95
addWeighted::F64ArithmTest::(1920x1080, 64FC1) 10.255 9.885 1.04
compare::F64ArithmTest::(127x61, 64FC1) 0.099 0.035 2.85
compare::F64ArithmTest::(640x480, 64FC1) 3.834 1.300 2.95
compare::F64ArithmTest::(1280x720, 64FC1) 11.456 3.847 2.98
compare::F64ArithmTest::(1920x1080, 64FC1) 25.670 8.738 2.94
divide::F64ArithmTest::(127x61, 64FC1) 0.064 0.065 0.98
divide::F64ArithmTest::(640x480, 64FC1) 2.493 2.352 1.06
divide::F64ArithmTest::(1280x720, 64FC1) 7.359 7.192 1.02
divide::F64ArithmTest::(1920x1080, 64FC1) 16.306 15.548 1.05
max::F64ArithmTest::(127x61, 64FC1) 0.029 0.026 1.12
max::F64ArithmTest::(640x480, 64FC1) 1.583 1.600 0.99
max::F64ArithmTest::(1280x720, 64FC1) 4.746 4.963 0.96
max::F64ArithmTest::(1920x1080, 64FC1) 9.827 10.143 0.97
min::F64ArithmTest::(127x61, 64FC1) 0.029 0.026 1.10
min::F64ArithmTest::(640x480, 64FC1) 1.625 1.630 1.00
min::F64ArithmTest::(1280x720, 64FC1) 4.821 5.000 0.96
min::F64ArithmTest::(1920x1080, 64FC1) 9.363 11.126 0.84
multiply::F64ArithmTest::(127x61, 64FC1) 0.029 0.027 1.10
multiply::F64ArithmTest::(640x480, 64FC1) 1.680 1.762 0.95
multiply::F64ArithmTest::(1280x720, 64FC1) 4.835 4.929 0.98
multiply::F64ArithmTest::(1920x1080, 64FC1) 10.427 10.564 0.99
subtract::F64ArithmTest::(127x61, 64FC1) 0.027 0.026 1.06
subtract::F64ArithmTest::(640x480, 64FC1) 1.605 1.660 0.97
subtract::F64ArithmTest::(1280x720, 64FC1) 4.758 5.016 0.95
subtract::F64ArithmTest::(1920x1080, 64FC1) 11.156 10.594 1.05
```
x-factor geomean **1.13x**. `compare` is a robust ~2.9x; most other ops land near parity because GCC already auto-vectorizes the simple scalar f64 loops at `-O3` on RVV (exactly the issue #30066 works around with `-fno-tree-vectorize`). Per-case values near parity swing with the known K1 bimodal timing behavior.
With `-fno-tree-vectorize` added to the build flags (the direction proposed in #30066 for RVV), the same A/B shows the kernel-level gains directly:
```
Min (ms)
Name of Test nt nt nt
base cand cand
min4 min4 min4
vs
nt
base
min4
(x-factor)
absdiff::F64ArithmTest::(127x61, 64FC1) 0.087 0.026 3.41
absdiff::F64ArithmTest::(640x480, 64FC1) 3.351 1.423 2.36
absdiff::F64ArithmTest::(1280x720, 64FC1) 10.189 4.595 2.22
absdiff::F64ArithmTest::(1920x1080, 64FC1) 22.476 10.094 2.23
add::F64ArithmTest::(127x61, 64FC1) 0.037 0.027 1.35
add::F64ArithmTest::(640x480, 64FC1) 1.538 1.446 1.06
add::F64ArithmTest::(1280x720, 64FC1) 4.644 4.403 1.05
add::F64ArithmTest::(1920x1080, 64FC1) 10.316 9.629 1.07
addWeighted::F64ArithmTest::(127x61, 64FC1) 0.076 0.027 2.81
addWeighted::F64ArithmTest::(640x480, 64FC1) 2.978 1.363 2.18
addWeighted::F64ArithmTest::(1280x720, 64FC1) 9.141 4.118 2.22
addWeighted::F64ArithmTest::(1920x1080, 64FC1) 19.856 9.075 2.19
compare::F64ArithmTest::(127x61, 64FC1) 0.103 0.031 3.31
compare::F64ArithmTest::(640x480, 64FC1) 3.993 1.225 3.26
compare::F64ArithmTest::(1280x720, 64FC1) 11.801 3.596 3.28
compare::F64ArithmTest::(1920x1080, 64FC1) 26.611 8.125 3.28
divide::F64ArithmTest::(127x61, 64FC1) 0.156 0.065 2.41
divide::F64ArithmTest::(640x480, 64FC1) 6.095 2.330 2.62
divide::F64ArithmTest::(1280x720, 64FC1) 18.191 6.880 2.64
divide::F64ArithmTest::(1920x1080, 64FC1) 40.743 15.331 2.66
max::F64ArithmTest::(127x61, 64FC1) 0.085 0.025 3.36
max::F64ArithmTest::(640x480, 64FC1) 3.272 1.428 2.29
max::F64ArithmTest::(1280x720, 64FC1) 10.009 4.581 2.18
max::F64ArithmTest::(1920x1080, 64FC1) 21.962 10.130 2.17
min::F64ArithmTest::(127x61, 64FC1) 0.086 0.025 3.40
min::F64ArithmTest::(640x480, 64FC1) 3.316 1.417 2.34
min::F64ArithmTest::(1280x720, 64FC1) 10.133 4.583 2.21
min::F64ArithmTest::(1920x1080, 64FC1) 22.334 10.160 2.20
multiply::F64ArithmTest::(127x61, 64FC1) 0.062 0.024 2.61
multiply::F64ArithmTest::(640x480, 64FC1) 2.430 1.362 1.78
multiply::F64ArithmTest::(1280x720, 64FC1) 7.427 4.344 1.71
multiply::F64ArithmTest::(1920x1080, 64FC1) 16.160 9.370 1.72
subtract::F64ArithmTest::(127x61, 64FC1) 0.038 0.025 1.48
subtract::F64ArithmTest::(640x480, 64FC1) 1.696 1.523 1.11
subtract::F64ArithmTest::(1280x720, 64FC1) 4.794 4.805 1.00
subtract::F64ArithmTest::(1920x1080, 64FC1) 12.058 10.820 1.11
```
x-factor geomean **2.09x** (36 cases, 1.00-3.41x): compare ~3.3x, divide ~2.6x, min/max ~2.2x, absdiff ~2.2x, addWeighted ~2.2x, multiply ~1.7-2.6x. `add`/`subtract` stay near parity in both regimes — they appear to be served by a different (already-vectorized) path before the kernel table is consulted.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
core: RVV HAL transpose2d for 3-channel element sizes - #30075
### Summary
`cv::transpose` for 3-channel types (`8UC3`/`8SC3`, `16UC3`/`16SC3`/`16FC3`, `32SC3`/`32FC3`) falls back to the scalar core path on RISC-V. The RVV HAL `transpose2d` dispatches on *element size* and only implements `esz ∈ {1, 2, 4, 8}`. This PR adds kernels for
`esz = 3, 6, 12`, and an accuracy test that covers the new paths including non-continuous ROI.
### Root cause
- `hal/riscv-rvv/src/core/transpose.cpp`: the `Transpose2dFunc tab[]` table has entries only at indices 1/2/4/8, so `cv_hal_transpose2d` returns `CV_HAL_ERROR_NOT_IMPLEMENTED` for every 3-channel type and the caller runs the generic path.
- On the scalable RVV backend (`intrin_rvv_scalable.hpp`) `CV_SIMD128` is never defined, so the existing `transpose_8/16/32/48bit_simd` fast paths `modules/core/src/matrix_transform.cpp` are compiled out entirely; what runs is the scalar 4x4 element loop. 8UC3 (the default color type) is the most visible case.
### Implementation
The new kernels deinterleave the three channels of several source rows with `vlseg3e{L}` and emit the transposed rows with strided segment stores (`vssseg{6,8}e{L}`). The number of source rows per block was picked experimentally: 8 rows for 8-bit lanes, 2 rows for 16/32-bit lanes. For 16/32-bit lanes, taller blocks e.g. 8 rows will add register pressure and cause regression. A one-row tail covers the remainder; unaligned 16/32-bit inputs fall back to the generic path.
### Performance
SpacemiT K1 / X60, OpenCV 5.x, `--perf_force_samples=20
--perf_min_samples=20`, `BinaryOpTest.transpose2d`:
| type (esz) | 640x480 | 1280x720 | 1920x1080 |
|---|---:|---:|---:|
| `CV_8UC3` (3) | 2.758 -> 1.648 ms (**1.67x**) | 16.772 -> 10.575 ms (**1.59x**) | 42.901 -> 39.552 ms (1.09x) |
| `CV_16SC3` (6) | 8.384 -> 3.556 ms (**2.36x**) | 31.112 -> 19.450 ms (**1.60x**) | 71.921 -> 62.460 ms (1.15x) |
Untouched element sizes (esz 1/2/4/8, 16 cases at 640x480 + 1280x720): geomean **1.005x**,
range 0.977x ... 1.058x.
### Accuracy
New test `Core_Transpose.C3ElementSizesWithRoi` (`modules/core/test/test_mat.cpp`) sweeps
`CV_8UC3, CV_8UC(6), CV_16SC3, CV_16FC3, CV_8UC(12), CV_32FC3` over sizes `1x1, 2x3, 137x5, 133x4` through non-continuous ROI views, byte-compares every element against the source, and asserts the bytes outside the destination ROI are untouched. On K1:
```
[ PASSED ] 4 tests. # Core_Transpose.C3ElementSizesWithRoi
# Core_Transpose/ElemWiseTest.accuracy/0
# Core_Rotate/ElemWiseTest.accuracy/0
```
`BinaryOpTest.transpose2d` already provides the performance coverage; no `opencv_extra` data is needed.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Carotene resize left edge clamp 5x - #30076
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [ ] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
dnn: reduce coordinate work in CPU Deconvolution - #30001
CPU `Deconvolution` computes coordinates and checks kernel bounds for each output value. This change builds valid tap tables for each axis, reuses outer coordinates per row, and uses SIMD for consecutive inputs at horizontal stride 1 or 2. It computes input addresses once per run and reuses them for each group of four outputs. Other positions use the scalar loop. Each output keeps the same sum order. The 1×1 path adds bias one channel segment at a time.
The table compares `5.x` at `8d126c4a` with this patch at `d0d1825`. It measures layer `forward()`, including GEMM, table setup, output reconstruction, and bias. The host is a Core Ultra 7 255H, with GCC 13.3, Release, one thread, and WSL2. IPP and OpenCL are disabled.
| Input, channels in/out | Kernel, stride | Upstream | PR | Speedup |
| --- | --- | ---: | ---: | ---: |
| 128×128, 16/16 | 1×1, 1 | 3.023 ms | 0.089 ms | 34.2× |
| 128×128, 16/16 | 2×2, 2 | 76.145 ms | 0.489 ms | 155.9× |
| 64×64, 16/16 | 3×3, 1, pad 1 | 10.709 ms | 0.239 ms | 44.8× |
| 16×16×16, 8/8 | 3×3×3, 1 | 21.843 ms | 0.472 ms | 46.3× |
Each time is the median of seven process medians, with two warmup calls and 31 measured calls per process. Version order alternates. Speedup uses the unrounded times.
Each worker reuses its address lists across rows. Their measured buffer capacities are 16 bytes for the 2×2 case and 96 bytes for the 3×3 case, excluding vector objects and allocator overhead. The SIMD run table is shared by the workers and uses one `size_t` per output column.
All 56 accuracy tests (8 new, 48 existing) pass with one and four threads. The eight performance tests pass. The 768 output comparisons match the base version. The extracted helper passes ASan/UBSan checks with SIMD and scalar code. The new accuracy tests use an independent scatter reference; duplicate pointwise and dynamic-weight cases were removed.
[Test code, raw results, and build steps](https://github.com/Hosi121/opencv/tree/bench/dnn-col2im-workspace/col2im-bench).
This changes the legacy CPU `Deconvolution` path. The new ONNX `ConvTranspose2` path is separate. These are layer results, not complete-model results. ARM hardware was not tested.
English is not my first language, so I used AI to help write this post.
Verify assembler can encode AVX-VNNI before enabling it - #29844
Closes: https://github.com/opencv/opencv/issues/29840
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
v_select had no 64-bit integer form on any backend except WASM, so code
that needs to blend v_int64/v_uint64 through a comparison mask had to
apply the mask by hand. intrin_sse.hpp carried the two entries commented
out as TBD; the other backends simply stopped at 32-bit and float64.
Every backend already blends bitwise or through a byte/boolean cast, so
the lane width does not change the operation:
SSE _mm_blendv_pd under CV_SSE4_1, the xor/and/xor form otherwise
AVX _mm256_blendv_epi8, as the narrower types already use
NEON vbslq_u64 / vbslq_s64
MSA msa_bslq_u8 over the byte reinterpretation
VSX vec_sel with the same boolean cast v_float64x2 uses
LSX/LASX __lsx_vbitsel_v / __lasx_xvbitsel_v on the raw register
RVV 0.7.1 vmerge_vvm with the b64 mask
RVV __riscv_vmerge, as for the narrower types
WASM already had both and is unchanged, and the generic v_reg
implementation in intrin_cpp.hpp already covers every type.
The 64-bit test chains could not simply call test_mask(), because its
v_signmask() expectations assume more than two lanes. Added test_select(),
which exercises v_select alone, and called it from TheTest<v_uint64>() and
TheTest<v_int64>().
imgproc: speed up getRectSubPix 8U->32F - #30050
### Summary
`getRectSubPix_8u32f`, the 8U-source -> 32F-patch path of `cv::getRectSubPix`, propagated the horizontal interpolation factor across the row:
```cpp
float prev = (1 - a)*(b1*src[0] + b2*src[src_step]);
for (int j = 0; j < win_size.width; j++) {
float t = a12*src[j+1] + a22*src[j+1+src_step];
dst[j] = prev + t;
prev = (float)(t*s); // loop-carried dependency
}
```
That recurrence is a serial dependency chain running the length of every row, so the loop neither pipelines nor vectorizes.
| patch size | 8U→32F | 32F→32F |
|---|---|---|
| 128×128 | 31.1 ms | 13.5 ms |
The recurrence is only a strength reduction. The bilinear weights are identical for every output pixel of a patch, so applying the constant 4-tap kernel directly gives the same result. This patch does that and vectorizes the loop with universal intrinsics (`vx_load_expand_q` → `v_cvt_f32` → 4× `v_mul`/`v_add` → `v_store`), keeping a scalar tail. `v_add`/`v_mul` are used instead of operators because the RVV scalable backend defines no operator overloads for its native vector types.
### Performance
SpacemiT K1 (RVV 1.0 / X60), `opencv_perf_imgproc --gtest_filter=*GetRectSubPix*`, `--perf_force_samples=20 --perf_min_samples=20`, `ALGO_HINT_ACCURATE`:
| patch | before | after | speedup |
|---|---|---|---|
| 16×16 | 0.610 ms | 0.308 ms | 1.98× |
| 32×32 | 2.116 ms | 0.843 ms | 2.51× |
| 64×64 | 8.016 ms | 2.822 ms | 2.84× |
| 128×128 | 31.109 ms | 10.369 ms | 3.00× |
The 8U→8U and 32F→32F paths are unchanged (0.99×–1.01×), and the whole `GetRectSubPix` perf suite improves from 81.66 ms to 53.95 ms (1.51×).
### Accuracy
`opencv_test_imgproc --gtest_filter=*GetRectSubPix*` passes on the same build.
```
Note: Google Test filter = *GetRectSubPix*
[==========] Running 1 test from 1 test case.
[----------] Global test environment set-up.
[----------] 1 test from Imgproc_GetRectSubPix
[ RUN ] Imgproc_GetRectSubPix.accuracy
[ OK ] Imgproc_GetRectSubPix.accuracy (945 ms)
[----------] 1 test from Imgproc_GetRectSubPix (945 ms total)
[----------] Global test environment tear-down
[==========] 1 test from 1 test case ran. (945 ms total)
[ PASSED ] 1 test.
```
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
<!-- Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the PR title. -->
Switches the warning for unknown or unsupported picture formats in
cap_ffmpeg_impl.hpp from CV_LOG_WARNING to CV_LOG_ONCE_WARNING.
This avoids log spam on every decoded frame when reading video
streams with CAP_PROP_CONVERT_RGB=0.
Generalize tokenizer loading to support method-based family dispatch in DNN - #29675
Companion PR: https://github.com/opencv/opencv_extra/pull/1402
### Changes
Added ALBERT and BERT support end-to-end, with `samples/dnn/albert_inference.py `and `samples/dnn/bert_inference.py `as validation samples, plus expanded coverage in modules/dnn/test/test_tokenizer.cpp. Required changes to `cv::dnn::dnn.hpp`, `graph_fusion_attention.cpp`, and `unicode.cpp/unicode.hpp` to support Unigram and WordPiece tokenizers.
To back ALBERT/BERT, generalized `cv::dnn::Tokenizer `from a single BPE implementation into a method-dispatched frontend, adding `core_wordpiece.cpp/hpp` (WordPiece) and `core_unigram.cpp/hpp` (Unigram) as new backends. `tokenizer.cpp` now routes by method across BPE, Gemma,, SentencePiece, Unigram, and WordPiece behind one shared interface.
Tested against the following samples and the output matches to old tokenizer:
```
gpt2_inference.py
qwen_inference.py
gemma3_inference.py
```
GPT2:
```
Preparing GPT-2 model...
Inferencing GPT-2 model...
Hello, I'm a language model, not a programming language. I'm a language model. I'm a language model. I'm a language model. I'm a language model. I'm a
```
Gemma3:
```
Preparing Gemma3 model...
Prompt:
<start_of_turn>user
What is OpenCV?<end_of_turn>
<start_of_turn>model
Inferencing Gemma3 model...
Response:
Okay, let's break down what OpenCV is.
**What is OpenCV?**
OpenCV (Open Source Computer Vision Library) is a powerful and
```
Qwen2.5:
```
Preparing Qwen2.5 model...
Prompt:
<|im_start|>user
What is OpenCV?<|im_end|>
<|im_start|>assistant
Inferencing Qwen2.5 model...
Response:
OpenCV is a set of computer vision libraries in C++ designed to be used for image and video processing. It provides a wide range of tools and functions for
```
### Tokenizer References
Byte-level BPE: [tokenizers/src/pre_tokenizers/byte_level.rs](https://github.com/huggingface/tokenizers/blob/main/tokenizers/src/pre_tokenizers/byte_level.rs)
(This defines the byte-level mapping rules, which is used in conjunction with the [BPE model](https://www.google.com/search?q=https://github.com/huggingface/tokenizers/blob/main/tokenizers/src/models/bpe/mod.rs))
SentencePiece BPE (Metaspace): [tokenizers/src/pre_tokenizers/metaspace.rs](https://www.google.com/search?q=https://github.com/huggingface/tokenizers/blob/main/tokenizers/src/pre_tokenizers/metaspace.rs)
(This defines the rule for replacing whitespace with the U+2581 _ character and handling byte fallback)
Unigram: [tokenizers/src/models/unigram/mod.rs](https://www.google.com/search?q=https://github.com/huggingface/tokenizers/blob/main/tokenizers/src/models/unigram/mod.rs)
(This contains the core logic for the Unigram lattice scoring and probabilistic tokenization rules)
WordPiece: [tokenizers/src/models/wordpiece/mod.rs](https://www.google.com/search?q=https://github.com/huggingface/tokenizers/blob/main/tokenizers/src/models/wordpiece/mod.rs)
(This explicitly cites Schuster & Nakajima in the code comments and implements the greedy longest-match rule with the ## prefix)
### Pull Request Readiness Checklist
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on
code under GPL or another license incompatible with OpenCV.
- [x] The PR is proposed to the proper branch (`5.x`).
- [x] There is a reference to the original bug report and related work.
- [x] There is accuracy test and test data in `opencv_extra`, same
branch name (`generalized-tokenizer`) — `bert/`, `t5/` fixtures
back the new C++ tests.
- [x] The feature is documented and sample code builds with project CMake.
Adds COLOR_BGR2Oklab, COLOR_RGB2Oklab, COLOR_Oklab2BGR, and
COLOR_Oklab2RGB to cvtColor, implementing Bjorn Ottosson's Oklab
perceptual color space (https://bottosson.github.io/posts/oklab/),
requested in #21760.
Scope for this first PR: CV_8U and CV_32F only (matching the existing
Lab/Luv precedent -- neither supports CV_16U either, likely for the
same reason: the cube-root nonlinearity doesn't lend itself to a
fixed-point integer fast path), plain scalar C++ (no SIMD, no OpenCL
kernel -- cvtColor's CV_OCL_RUN/CV_Assert machinery already falls back
to this CPU path cleanly for codes without an OpenCL case), and the
core RGB<->Oklab conversion only (not the Okhsv/Okhsl variants the
issue also mentions, which the issue itself calls secondary to the
"most directly important" Oklab conversion).
Encoding: CV_32F keeps Oklab's native range (L in [0,1], a/b roughly
in [-0.5,0.5] for in-gamut sRGB colors), matching how the reference
formulas and other tools (e.g. CSS Color 4's oklab()) represent it.
CV_8U scales L by 255 and offsets a/b by 128, mirroring CIE Lab's own
8-bit encoding shape.
Testing:
- The forward matrices were verified against an independent NumPy
reimplementation before writing any C++: round-trip identity over
20000 random RGB samples (~1.6e-6 max error), plus cross-checked
pure red/green/blue/white/black/gray against commonly-published
Oklab reference values -- matches to 6+ significant figures.
- New tests in test_color.cpp: reference-value cross-check (same
known values as above, now through the actual cv::cvtColor), a
channel-order check (COLOR_RGB2Oklab vs COLOR_BGR2Oklab on the same
logical pixels), float32 and uint8 round-trips, and alpha-channel
passthrough on the 4-channel inverse. All pass; verified the uint8
round-trip test actually exercises meaningful precision by comparing
its tolerance against CIE Lab's own established uint8 tolerance
(CV_ColorLabTest accepts up to 37 for the analogous case; this uses
32 for a still-meaningful, not looser, bound).
- Confirmed swapBlue() needed COLOR_BGR2Oklab/COLOR_Oklab2BGR added to
its explicit list -- its default (true) is correct for the RGB
variants by accident but wrong for the BGR ones, which would have
silently swapped red and blue. Caught by code review before testing,
then confirmed fixed via the channel-order test.
- Ran the full opencv_test_imgproc suite: no new failures (the same
360 pre-existing "missing opencv_extra test data" failures present
before this change, e.g. StackBlur/HoughCircles/ColorBayer, all in
unrelated files).
- Verified the UMat (OpenCL) path falls back to the CPU implementation
correctly (no OpenCL kernel is added in this PR).
- Added a perf test (cvtColorOklab) covering both directions, CV_8U
and CV_32F, at VGA/720p/1080p.
Natural follow-ups, intentionally left out of this first PR: SIMD/
fixed-point optimization (this is a plain scalar implementation),
Okhsv/Okhsl, and an OpenCL kernel.
Fix function pointer signature mismatches - #30040
Contrib PR: https://github.com/opencv/opencv_contrib/pull/4224
Calling a function through a function pointer with a mismatched signature is undefined behavior in C/C++ and causes Clang Control Flow Integrity to trap with `SIGILL` (`ud1`) at indirect call sites.
This is a follow-up on https://github.com/opencv/opencv/pull/28939
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake