26154 Commits
Author SHA1 Message Date
Alexander Smorkalov bed6800f05 Merge pull request #29940 from sergiud:fisheye-lm-solver
calib3d: improve fisheye extrinsic refinement
2026-09-24 09:55:32 +03:00
Pratham Kumar f46dd2af02 Merge pull request #30045 from pratham-mcw:svd_opt
core: add ArmPL HAL backend for cv::SVD::compute - #30045

## Summary

- Adds an ARM Performance Library (ARMPL) HAL backend for cv::SVD::compute (cv_hal_SVD32f / cv_hal_SVD64f), replacing OpenCV's default SVD path with ARMPL's LAPACKE_sgesdd / LAPACKE_dgesdd for float and double matrices on AArch64. 

## Changes
**hal/armpl/include/armpl_hal_core.hpp:**

- Declare armpl_hal_SVD32f / armpl_hal_SVD64f and register them as cv_hal_SVD32f / cv_hal_SVD64f.

**hal/armpl/src/armpl_hal_core.cpp:**

- Add a templated armpl_svd helper that maps OpenCV's CV_HAL_SVD_NO_UV / SHORT_UV / MODIFY_A / FULL_UV flags onto LAPACKE's jobz parameter and calls LAPACKE_sgesdd / dgesdd.
- Fall back to OpenCV's default (non-HAL) implementation for small matrices (m < 33, ARMPL_SVD_SMALL_MATRIX_THRESH).
## Performance Benchmarks
<img width="842" height="422" alt="image" src="https://github.com/user-attachments/assets/12b84943-be57-4bf4-9d7e-c85121da4689" />
2026-09-24 09:22:41 +03:00
Loic GoossensandClaude Opus 5 527a6b2284 videoio(ffmpeg): guard seek() against AV_NOPTS_VALUE start_time
CvCapture_FFMPEG::seek() seeds the seek target with the stream start_time
without checking for AV_NOPTS_VALUE, while dts_to_sec() right above it does
perform that check.

Some containers do not let FFmpeg establish a start time, for instance
Matroska files whose H.264 track is declared through the legacy VfW wrapper
(CodecID V_MS/VFW/FOURCC with a BITMAPINFOHEADER) instead of V_MPEG4/ISO/AVC
with an avcC record. For those, start_time is AV_NOPTS_VALUE, the computed
target becomes INT64_MIN plus an offset, and av_seek_frame() with
AVSEEK_FLAG_BACKWARD lands at position 0. The refinement loop below then
decodes every single frame up to the requested one, so every seek degrades to
a linear scan. The returned frame is still correct, which makes the failure
silent and easy to miss.

Such a container is not otherwise malformed: its Cues index is complete and
correct, and `ffmpeg -ss` seeks it in constant time at any offset.

Measured with OpenCV 4.14.0 on a 15 h H.264 recording (1628443 frames),
Windows x64, stock prebuilt FFmpeg wrapper versus the same wrapper rebuilt
with this patch:

  target        before      after
     10 s      0.246 s     0.092 s
     60 s      1.494 s     0.039 s
    300 s      7.569 s     0.052 s
  10000 s      ~4 min      0.061 s
  54000 s     ~22 min      0.051 s

Before the patch the cost grew strictly linearly with the target position, at
roughly 25 ms per second of video, so the two longest targets were
extrapolated rather than waited out. Decoded frames are unchanged for targets
that both paths reach.

The guard mirrors the one already present in dts_to_sec().

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 10:05:58 +02:00
Sewon Ahn 62751db0d7 Merge pull request #30005 from lrycro:fix/glob-readdir-leak
core: fix memory leak in glob()'s readdir() on WinRT/_WIN32_WCE - #30005

## Problem

Fixes #30004

In `modules/core/src/glob.cpp`, the WinRT/`_WIN32_WCE` implementation of `readdir()` allocates a new buffer for `dir->ent.d_name` on every call and overwrites the previous pointer without freeing it:

```cpp
char* aname = new char[asize+1];
...
dir->ent.d_name = aname;
```

`cv::glob()` calls `readdir()` once per directory entry, so every call except the last leaks its allocation. Additionally, the `DIR` destructor that releases `d_name` was gated by `#ifdef WINRT` only, so `_WIN32_WCE` builds leaked every allocation, including the last one.

## Fix

- Free the previous `dir->ent.d_name` before overwriting it in `readdir()`, matching how `~DIR()` already frees it on WinRT.
- Extend the `DIR` destructor guard from `#ifdef WINRT` to `#if defined(WINRT) || defined(_WIN32_WCE)` so the final buffer is also released under `_WIN32_WCE`.

## Checklist

- [x] I agree to contribute to the project under the Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on code under GPL or another license incompatible with OpenCV.
- [x] The PR is proposed to the proper branch (`4.x`).
- [x] There is a reference to the original bug report and related work.
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable.
- [x] The feature is well documented and sample code can be built with the project CMake.

Thanks for reviewing.
2026-09-21 18:07:13 +03:00
pbkx 018bc3ebdf fix OutputArray release for std::array<Mat> 2026-09-20 20:35:14 -07:00
Sergiu Deitsch 9e15eac47f calib3d: improve fisheye extrinsic refinement
Improve fisheye extrinsic refinement for ill-conditioned views and
poorly scaled object points. Replace the Gauss-Newton update with a
Levenberg-Marquardt solver and normalize translation parameters to make
convergence less sensitive to problem scale.
2026-09-20 01:45:44 +02:00
yqtian-se ffaa389415 Scope GCC diagnostic suppression in RVV intrinsics 2026-09-18 12:37:29 +00:00
Alexander Smorkalov cf66d2c6ad Merge pull request #29972 from andreyfe1:extend_perf_calib3d
Extend calib3d perf tests
2026-09-17 16:06:48 +03:00
Alexander Smorkalov 7434859646 Merge pull request #29960 from amd:imp_mergedebevec
Fuse and vectorize MergeDebevec for HDR.
2026-09-17 15:59:38 +03:00
Madan mohan Manokar 46b9862506 Fuse and vectorize MergeDebevec for 1- and 3-channel HDR.
Fold LUT/split/mul into one pass and SIMD the accumulate/finalize loops to cut full-image temporaries.
2026-09-17 09:21:54 +00:00
Fedorov, Andrey 71c17ec0cd calib3d perf: filterSpeckles 2026-09-17 02:00:34 -07:00
pbkx db189ed863 fix calcCovarMatrix vector mean ROI handling 2026-09-16 18:41:52 -07:00
Abhishek Gola 5c061f003b Merge pull request #29312 from pratham-mcw/dft_cache_thrashing_fix
core: Add padded buffer for DFT on Windows ARM64 to reduce cache thrashing
2026-09-15 18:07:26 +05:30
Alexander Smorkalov 62f14c08d1 Merge pull request #29956 from asmorkalov:as/win32_warning_fix
Warning fix on Windows.
2026-09-15 12:01:27 +03:00
Alexander Smorkalov b24763ffa9 Warning fix on Windows. 2026-09-15 10:21:31 +03:00
Andrei Fedorov 96ec0c97df Merge pull request #29922 from andreyfe1:extend_perf_core
Extend core performance tests - #29922

Added more cases for performance tests. It's a part of a bigger PR https://github.com/opencv/opencv/pull/29631
execution time increased by <18% per my measurements.
Duplication of https://github.com/opencv/opencv/pull/29900, but from a different fork.
`opencv_extra` PR is https://github.com/opencv/opencv_extra/pull/1410

Following was added:

| File | Test | Added |
|------|------|-------|
| `perf_arithm.cpp` | `BinaryOpTest.*` (add/subtract/multiply/absdiff/min/max/transpose2d) | types `CV_16UC1`, `CV_64FC1` |
| `perf_compare.cpp` | `compareScalar` | types `CV_16UC1`, `CV_16SC1` |
| `perf_dot.cpp` | `dot` | type `CV_64FC1` |
| `perf_flip.cpp` | `flip` (`FLIP_TYPES`) | types `CV_32FC3`, `CV_32FC4` |
| `perf_mat.cpp` | `Mat_CopyToWithMask` | types `CV_32FC3` |
| `perf_mat.cpp` | `Mat_SetToWithMask` | types `CV_8UC3, CV_8UC4, CV_16UC3, CV_16UC4, CV_32FC3, CV_32FC4` |
| `perf_norm.cpp` | `norm` | norm type `NORM_L2SQR` |
| `perf_norm.cpp` | `norm_mask` | types `CV_8UC3`, `CV_16UC3`, `CV_32FC3`; norm type `NORM_L2SQR` |
| `perf_norm.cpp` | `norm2` | norm types `NORM_L2SQR`, `NORM_RELATIVE+NORM_L2SQR` |
| `perf_norm.cpp` | `norm2_mask` | types `CV_8UC3`, `CV_16UC3`, `CV_32FC3`; norm types `NORM_L2SQR`, `NORM_RELATIVE\|NORM_L2SQR` |
| `perf_sort.cpp` | `sort`, `sorIdx` (`TYPICAL_MAT_TYPES_SORT`) | types `CV_16SC1`, `CV_32SC1`, `CV_64FC1` |
| `perf_stat.cpp` | `sum`, `mean` | types widened `{8UC1,8UC4,32FC1}` → `8U/16U/16S/32F × C1/C3/C4` |
| `perf_stat.cpp` | `mean_mask`, `meanStdDev`, `meanStdDev_mask` | types widened `{8UC1,8UC4,32FC1}` → `8U/16U/32F × C1/C3/C4` |

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [ ] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
2026-09-15 08:47:00 +03:00
Alexander Smorkalov da8527c189 Merge pull request #29862 from kjg0724:variational-refinement-simd
video: migrate VariationalRefinement SIMD kernels to universal/scalable intrinsics
2026-09-14 16:51:16 +03:00
Alexander Smorkalov cd3aa54bc7 Merge pull request #29943 from sergiud:rational-model-docs
calib3d: clarify rational distortion model
2026-09-14 15:13:56 +03:00
Minsoo Kim 329e9c279c Merge pull request #29913 from 33modeling:docs-filter2d-integer-kernels-22243
imgproc: document integer filter2D kernels 🤖🤖🤖 - #29913

### Summary

`filter2D` accepts integer kernels, but its parameter documentation currently restricts kernels to floating-point matrices. Document the existing integer support and add parameterized accuracy coverage against `cvtest::filter2D`, the reference implementation.

The 40 cases cover five integer kernel depths (`CV_8U`, `CV_8S`, `CV_16U`, `CV_16S`, `CV_32S`), 3x3 and 13x13 kernels, and four source formats (`CV_8UC1`, `CV_32FC1`, `CV_32FC3`, `CV_64FC1`). Signed kernels include negative coefficients. Kernel depth, size, and source format are separate Google Test parameters.

Fixes #22243. Related work: #29607 was closed without merging. This follow-up accounts for its review requests about parameterized tests and avoids claiming a specific implementation path from a test name. No filtering implementation changes are included.

### Validation

Built `opencv_test_imgproc` from current 4.x (`bb9e3eff13`) on Linux with GCC 13.3.0, Release/SSE3. IPP, OpenCL, and additional CPU dispatch variants were disabled in this local build.

```
./bin/opencv_test_imgproc --gtest_filter='*Imgproc_Filter2D*:*Imgproc_FilterSupportedFormats*'
```

All 58 selected tests passed: 40 new cases and 18 existing filter accuracy/format/regression cases. `git diff --check` passed. This is targeted CPU validation, not a full imgproc suite or a GPU/backend validation claim. The tests use synthetic data and need no opencv_extra assets.

### Pull Request Readiness Checklist

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on code under GPL or another license incompatible with OpenCV.
- [x] The PR is proposed to the proper branch: 4.x for a documentation correction applicable to both maintained branches.
- [x] There is a reference to the original bug report and related work.
- [x] Accuracy tests are included; no external test data or performance change is involved.
- [x] The parameter documentation is updated and the affected test target builds.

### AI assistance

OpenAI Codex assisted with the source/review audit, patch preparation, and local validation. The commit includes an Assisted-by trailer.
2026-09-14 11:51:24 +03:00
Sergiu Deitsch 87eff0a6d2 calib3d: clarify rational distortion model
Distinguish OpenCV's rational distortion model from the Rational
Function Model in the documentation.
2026-09-13 22:03:14 +02:00
Alexander Smorkalov 225283da63 Merge pull request #29927 from asmorkalov:as/backport_patchNans
PatchNans for double in 4.x
2026-09-11 20:49:54 +03:00
Alexander Smorkalov c1ca7c2e56 Merge pull request #29879 from spmallick:fix/imgcodecs-file-close-errors
imgcodecs: report JPEG and PNG output file errors 🤖🤖🤖
2026-09-11 20:48:52 +03:00
Alexander Smorkalov ec8d2e3810 Use default per strategy for optical flow tests. 2026-09-11 16:32:10 +03:00
Alexander Smorkalov 4ecb9762d1 PatchNans for double in 4.x 2026-09-11 15:33:26 +03:00
Alexander Smorkalov b808ef6615 Dropped AI generated test for disk io issues. 2026-09-11 11:40:25 +03:00
pbkx 32080eb949 Merge pull request #29894 from pbkx:fix-sunras-encoding-checks
imgcodecs: fix Sun Raster encoding checks 🤖🤖🤖 - #29894

SunRasterDecoder saves the header's ras_type field in m_encoding but several validation and decoding checks compared RAS_BYTE_ENCODED and RAS_FORMAT_RGB against m_type. This caused valid byte-encoded and RGB-format Sun Raster images to be rejected before their existing decoding paths could be used.

This patch uses the file encoding to make these decisions and limits byte encoding to 8-bit input and RGB-format input to the supported 24- and 32-bit paths. It also selects channel conversion from the file and requested output orders and avoids indexed-palette conversion when decoding truecolor inputs as grayscale.

The regression tests use Sun Raster files written by Netpbm's `pnmtorast -rle` and ImageMagick's SUN encoder. The matching `opencv_extra` PR (opencv/opencv_extra#1408) contains these files and lets them be inspected independently with compatible image viewers. The tests verify exact RLE literal and run values, truncated RLE input, 24- and 32-bit RGB-format input, BGR and RGB output, and grayscale conversion.

The Netpbm file also exposed an existing row-padding error in the RLE path: the decoder consumed a padding byte after even-width rows, although those rows require no padding. The patch now consumes that byte only for odd-width rows.

No public API is changed.

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
2026-09-11 08:56:54 +03:00
superbigcup325 600bcf4b21 python: avoid empty LD_LIBRARY_PATH/PATH entries on import (4.x backport of #29905)
Backport of #29905 to the 4.x branch, applied verbatim: after this change
`modules/python/package/cv2/__init__.py` is byte-identical to the 5.x version
(same blob 99685b3791), since both branches carried the same code here.

`import cv2` unconditionally rewrites the environment; when LD_LIBRARY_PATH
was previously unset this creates a dangling separator, i.e. an empty entry
which, per ld.so(8), resolves to the current working directory of every
subsequently spawned child process. Children may then pick up same-named
shared libraries from the CWD and fail (real-world case reported in
opencv/opencv-python#1268: a Nuitka-standalone app's `xdg-open -> kde-open`
dying with `libssl.so.3: version 'OPENSSL_3.2.0' not found`; reproduced on
4.11.0.86 - 4.14.0.94 manylinux wheels). The Windows PATH line has the same
dangling-separator pattern.

Identical to #29905: factor the prepend into _prepend_env_paths(), filter
empty entries, write the variable only when there is something to add, and
join with the old value without producing an empty entry.

The follow-up question of not mutating the process environment at all
remains tracked in #28994.
2026-09-10 23:52:36 +08:00
Alexander Smorkalov de42350f04 Disabled perf test for IntelligentScissors as it's too long. 2026-09-10 09:42:45 +03:00
Alexander Smorkalov 1c312e0820 Merge pull request #29908 from andreyfe1:andreyfe1/extend_perf_imgproc
Extend OpenCV imgproc performance tests
2026-09-10 09:04:40 +03:00
Alexander Smorkalov bb9e3eff13 Merge pull request #29858 from ahmadmasood43:boundingrect-29837-4x
imgproc: saturate out-of-range float coordinates in boundingRect
2026-09-09 15:14:10 +03:00
Fedorov, Andrey ed65b9186e changed imgproc perf tests 2026-09-09 04:34:34 -07:00
Sridhar 2d781f51e1 Merge pull request #29878 from sridhar-git05:fix-broadcast-zero-dimension
core: handle zero-sized broadcast dimensions #29878

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake

<!-- Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the PR title. -->

### Changes

- Added test coverage for the zero-sized dimension case in `cv::broadcast()`.
- The test exercises the `false` branch of `_flatten_for_broadcast()`.
- Fixed division by zero when the broadcast destination has zero elements.

### Test

- `opencv_test_core.exe --gtest_filter=BroadcastTo.*`

All `BroadcastTo` tests pass.

Fixes #28910
2026-09-09 12:15:52 +03:00
Abhishek Gola 42076ef3b7 Merge pull request #29326 from pratham-mcw/guided_filter-opt-simd
ximgproc: add NEON intrinsics support for BoxFilter Function
2026-09-08 16:58:17 +05:30
pbkx 7c3880d0de validate KeyPoint::convert indexes 2026-09-07 21:44:04 -07:00
CodeCraftsman 0c6b4a9b6d Merge pull request #29853 from Thebinary110:4x-fix-dis-opticalflow-overflow
video: fix DISOpticalFlow heap-buffer-overflow with patch_size > border_size - #29853

Fixes #20185. 4.x companion to #29715.

@asmorkalov asked to retarget #29715 to 4.x since the issue reproduces there too, but that PR's branch descends from 5.x, so literally changing its base produces an unreviewable ~2M-line diff (the two branches have diverged far beyond this module). Opening a separate PR instead: same two commits, cherry-picked cleanly onto 4.x's current tip with zero conflicts (`modules/video/src/dis_flow.cpp` is byte-for-byte identical between the branches apart from this fix).

### What / why
See #29715 for the full writeup. Summary: `DISOpticalFlowImpl` pads `I1` with a fixed 16px border, while `PatchInverseSearch`'s search-position clamp lets a patch be placed up to `patch_size - 1` px outside the image -- safe only while `patch_size <= border_size`. Since `patch_size` is user-settable with no upper bound relative to the hardcoded border, `setPatchSize()` past 16 (or a large-enough temporal-candidate flow) reads past the end of the padded buffer (confirmed via AddressSanitizer).

Fix (per review on #29715): rather than growing the persistent `border_size` object field to match `patch_size` (which a reviewer correctly flagged as the wrong place for a per-call derived quantity), the search clamp itself (`i/j_lower_limit`, `i/j_upper_limit` in `PatchInverseSearch_ParBody`) now accounts for the read window needing to stay inside the *existing* padded buffer. For `patch_size <= border_size` (every built-in preset) the bounds are algebraically identical to the originals -- `border_size` itself is untouched.

### Testing
Verified this reproduces identically on 4.x: built an ASan-instrumented Debug configuration (`core+imgproc+imgcodecs+features2d+flann+video+ts`), confirmed the crash reproduces on pristine 4.x (same crash site as the original report), confirmed it's gone with this fix, and ran the new `regression_20185_patch_larger_than_border` / `regression_20185_stress` tests (5 repeated runs) plus the full `opencv_test_video` suite -- no failures attributable to this change (everything else failing needs `opencv_extra` test data not configured in this scoped build).

### PR checklist
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on code under GPL or another incompatible license.
- [x] The PR is proposed to the proper branch (4.x, per maintainer request on #29715).
- [x] Accuracy tests included (see above).
- [x] No public API/behavior change, so no documentation or sample updates needed.
2026-09-07 21:43:55 +03:00
Satya Mallick 7f48fe9b49 imgcodecs: propagate JPEG and PNG output file errors 2026-09-06 17:07:33 +05:30
Alexander Smorkalov 415ba6a444 Merge pull request #29805 from amd:opencl_boxfilter
imgproc: Enable boxFilter OpenCL fast paths on Non-Intel GPUs.
2026-09-05 17:01:29 +03:00
ahmadmasood43 0b475cd66f imgproc: saturate out-of-range float coordinates in boundingRect
pointSetBoundingRect() passed CV_32F coordinates to cvFloor() without
honouring its documented INT_MIN..INT_MAX precondition, so a point set
outside the int range collapsed to a 1x1 rect at INT_MIN, while +inf and
-inf gave INT_MIN and INT_MAX - an inverted rectangle. The extrema are
now searched for in floats, as the vectorized path already did, and
floored and saturated to the closest representable bound once, after the
reduction.

The sides are computed in int64 and clamped to INT_MAX: xmax - xmin + 1
overflows int once the extrema saturate, and did so already for a CV_32S
point set spanning the whole int range.

Fixes #29837
2026-09-05 17:12:24 +05:00
Jeongkeun Kim 2d350df3c3 video: migrate VariationalRefinement SIMD kernels to universal/scalable intrinsics
- Type/stride substitution only; arithmetic and evaluation order unchanged.
- RedBlackSOR: the v_extract<3>(prev, next) previous-lane construction is
  replaced by an unaligned vx_load(p_next + j - 1); v_extract<3> is only
  previous-lane when vlanes == 4, so this is required for wider lanes.
- HorPass keeps a strict bound, j < len - vlanes, while the other three use
  j <= len - vlanes, since the vector body applies UPDATE to every lane and
  the rightmost element must not receive UPDATE across the right border.
- No SVE claim: there is no SVE backend in-tree; scalable means RVV today.
- RedBlackSOR: drop the pW_next_vec load, unused since the lane-shift now
  reads pW_next directly.
2026-09-04 09:50:10 +09:00
Madan mohan Manokar a15bcc7ea8 Merge pull request #29393 from amd:fast_moments
imgproc: Optimized Moments & perf test added - #29393

- Added dispatch for moments (CV_8U, CV_16U, CV_32F and CV_64F) 
- CV_16S is kept in scalar due to regresssions observed.
- Perf test added for meanShift, CamShift and matchShapes.
- Extend Moments1 perf coverage to CV_8U alongside 16/32/64-bit depths.

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
2026-09-03 16:38:40 +03:00
Alexander Smorkalov e39c9a2411 Merge pull request #29850 from lazerg:fix/ipc-float32-cps-underflow
imgproc: guard against zero magnitude in phaseCorrelateIterative
2026-09-03 14:35:25 +03:00
Alexander Smorkalov 1209808c70 Merge pull request #29855 from lazerg:fix/issue-29854-empty-input-crash
imgproc: reject empty input in connected components
2026-09-03 14:25:47 +03:00
Alexander Smorkalov 7b0ce2b50b Merge pull request #29848 from amd:imp_cvtcolor
imgproc: faster bit-exact BGR(A)2GRAY on x86 AVX-512
2026-09-03 12:35:13 +03:00
Lazizbek Ergashev e7652038ee imgproc: reject empty input in connected components 2026-09-02 23:11:05 +05:00
Lazizbek Ergashev bf4a8248c9 imgproc: guard against zero magnitude in phaseCorrelateIterative 2026-09-02 18:28:02 +05:00
Madan mohan Manokar 184fb3f51c imgproc: faster bit-exact BGR(A)2GRAY on x86 AVX-512
Use vpmaddwd on BGRA and vpermb+vpmaddwd on BGR; fall back to the universal SIMD kernel.
2026-09-02 10:31:18 +00:00
Alexander Smorkalov 2ce3cbc260 Merge pull request #29146 from XieFuwei:fix/sunras-grayscale-imread
imgcodecs: fix SunRaster IMREAD_GRAYSCALE returning black image
2026-09-02 13:25:22 +03:00
Pratham Kumar 8602d9d964 Merge pull request #29831 from pratham-mcw:nchw_nchw_opt
dnn: fix eltwise layer memory traversal pattern causing Windows-ARM64 slowdown - #29831

## Problem
- `EltwiseInvoker::operator()` (backs the `Eltwise` layer's SUM/PROD/DIV ops) ran slower on Windows-ARM64 than x64.

## Root cause
- The loop tiled the output into fixed-size blocks and looped over the **channel axis inside each block**.
- In NCHW layout, consecutive channels of the same sample sit far apart in memory (`planeSize * 4` bytes apart).
- So each block jumped between ~256 memory locations per tensor — roughly **768 concurrent strided streams** across the three tensors involved.
- Hardware prefetchers can only track a small, fixed number of streams and Windows-ARM64's prefetcher falls over on this pattern.

## Fix
- Walk the output buffer contiguously and track the current (sample, channel) plane from the offset and clamp each block so it never crosses into the next channel.
- This removes the inner channel loop entirely: one channel finishes before the next starts, so every tensor is read/written sequentially instead of jumping around.
- No behavior change, only the traversal order is different.

## Performance Benchmarks:
<img width="1479" height="208" alt="image" src="https://github.com/user-attachments/assets/0f878757-d988-4564-b166-96d80834c77e" />
2026-09-01 14:26:36 +03:00
Alexander Smorkalov c3565d2759 Merge pull request #29796 from asmorkalov:as/aravis_pixel_formats
Select pixel format in ARAVIS backend for VideoCapture
2026-09-01 10:36:05 +03:00
Pratham Kumar 6da81770bb Merge pull request #29815 from pratham-mcw:core-norm2_mask-simd-opt
core: vectorize masked norm/normDiff for remaining depths - #29815

### Summary

- The masked `cv::norm()` / `cv::norm(a, b)` kernels in `norm.simd.hpp` had SIMD specializations only for `uchar`, `ushort` and `float`. 
- `schar`, `short`, `int`, `double` and `uchar` L2 paths uses the scalar implementation.

### Changes

- Added new vectorized implementations of MaskedNorm{Inf,L1,L2}_SIMD for schar, short, int, and double.
- Added new vectorized implementation of MaskedNormL2_SIMD<uchar, int>.
- Added new MaskedNormDiff{Inf,L1,L2}_SIMD implementations for double.
- Added cn == 4 v_load_deinterleave paths to the uchar L1/L2 kernels.
- Replaced the single f64 accumulator in MaskedNormL1_SIMD<float,double> with four independent ones, so the widening adds can overlap instead of each waiting on the previous.

### Performance Benchmarks

<img width="715" height="709" alt="image" src="https://github.com/user-attachments/assets/46afea06-df98-4dd4-a346-dc8cd4a065c9" />
2026-09-01 10:35:00 +03:00