mirror of
https://github.com/opencv/opencv.git
synced 2026-10-05 04:03:32 +03:00
core: RVV HAL transpose2d for 3-channel element sizes - #30075 ### Summary `cv::transpose` for 3-channel types (`8UC3`/`8SC3`, `16UC3`/`16SC3`/`16FC3`, `32SC3`/`32FC3`) falls back to the scalar core path on RISC-V. The RVV HAL `transpose2d` dispatches on *element size* and only implements `esz ∈ {1, 2, 4, 8}`. This PR adds kernels for `esz = 3, 6, 12`, and an accuracy test that covers the new paths including non-continuous ROI. ### Root cause - `hal/riscv-rvv/src/core/transpose.cpp`: the `Transpose2dFunc tab[]` table has entries only at indices 1/2/4/8, so `cv_hal_transpose2d` returns `CV_HAL_ERROR_NOT_IMPLEMENTED` for every 3-channel type and the caller runs the generic path. - On the scalable RVV backend (`intrin_rvv_scalable.hpp`) `CV_SIMD128` is never defined, so the existing `transpose_8/16/32/48bit_simd` fast paths `modules/core/src/matrix_transform.cpp` are compiled out entirely; what runs is the scalar 4x4 element loop. 8UC3 (the default color type) is the most visible case. ### Implementation The new kernels deinterleave the three channels of several source rows with `vlseg3e{L}` and emit the transposed rows with strided segment stores (`vssseg{6,8}e{L}`). The number of source rows per block was picked experimentally: 8 rows for 8-bit lanes, 2 rows for 16/32-bit lanes. For 16/32-bit lanes, taller blocks e.g. 8 rows will add register pressure and cause regression. A one-row tail covers the remainder; unaligned 16/32-bit inputs fall back to the generic path. ### Performance SpacemiT K1 / X60, OpenCV 5.x, `--perf_force_samples=20 --perf_min_samples=20`, `BinaryOpTest.transpose2d`: | type (esz) | 640x480 | 1280x720 | 1920x1080 | |---|---:|---:|---:| | `CV_8UC3` (3) | 2.758 -> 1.648 ms (**1.67x**) | 16.772 -> 10.575 ms (**1.59x**) | 42.901 -> 39.552 ms (1.09x) | | `CV_16SC3` (6) | 8.384 -> 3.556 ms (**2.36x**) | 31.112 -> 19.450 ms (**1.60x**) | 71.921 -> 62.460 ms (1.15x) | Untouched element sizes (esz 1/2/4/8, 16 cases at 640x480 + 1280x720): geomean **1.005x**, range 0.977x ... 1.058x. ### Accuracy New test `Core_Transpose.C3ElementSizesWithRoi` (`modules/core/test/test_mat.cpp`) sweeps `CV_8UC3, CV_8UC(6), CV_16SC3, CV_16FC3, CV_8UC(12), CV_32FC3` over sizes `1x1, 2x3, 137x5, 133x4` through non-continuous ROI views, byte-compares every element against the source, and asserts the bytes outside the destination ROI are untouched. On K1: ``` [ PASSED ] 4 tests. # Core_Transpose.C3ElementSizesWithRoi # Core_Transpose/ElemWiseTest.accuracy/0 # Core_Rotate/ElemWiseTest.accuracy/0 ``` `BinaryOpTest.transpose2d` already provides the performance coverage; no `opencv_extra` data is needed. ### Pull Request Readiness Checklist See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request - [x] I agree to contribute to the project under Apache 2 License. - [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV - [x] The PR is proposed to the proper branch - [ ] There is a reference to the original bug report and related work - [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable Patch to opencv_extra has the same branch name. - [x] The feature is well documented and sample code can be built with the project CMake