CvCapture_FFMPEG::seek() seeds the seek target with the stream start_time
without checking for AV_NOPTS_VALUE, while dts_to_sec() right above it does
perform that check.
Some containers do not let FFmpeg establish a start time, for instance
Matroska files whose H.264 track is declared through the legacy VfW wrapper
(CodecID V_MS/VFW/FOURCC with a BITMAPINFOHEADER) instead of V_MPEG4/ISO/AVC
with an avcC record. For those, start_time is AV_NOPTS_VALUE, the computed
target becomes INT64_MIN plus an offset, and av_seek_frame() with
AVSEEK_FLAG_BACKWARD lands at position 0. The refinement loop below then
decodes every single frame up to the requested one, so every seek degrades to
a linear scan. The returned frame is still correct, which makes the failure
silent and easy to miss.
Such a container is not otherwise malformed: its Cues index is complete and
correct, and `ffmpeg -ss` seeks it in constant time at any offset.
Measured with OpenCV 4.14.0 on a 15 h H.264 recording (1628443 frames),
Windows x64, stock prebuilt FFmpeg wrapper versus the same wrapper rebuilt
with this patch:
target before after
10 s 0.246 s 0.092 s
60 s 1.494 s 0.039 s
300 s 7.569 s 0.052 s
10000 s ~4 min 0.061 s
54000 s ~22 min 0.051 s
Before the patch the cost grew strictly linearly with the target position, at
roughly 25 ms per second of video, so the two longest targets were
extrapolated rather than waited out. Decoded frames are unchanged for targets
that both paths reach.
The guard mirrors the one already present in dts_to_sec().
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
core: fix memory leak in glob()'s readdir() on WinRT/_WIN32_WCE - #30005
## Problem
Fixes#30004
In `modules/core/src/glob.cpp`, the WinRT/`_WIN32_WCE` implementation of `readdir()` allocates a new buffer for `dir->ent.d_name` on every call and overwrites the previous pointer without freeing it:
```cpp
char* aname = new char[asize+1];
...
dir->ent.d_name = aname;
```
`cv::glob()` calls `readdir()` once per directory entry, so every call except the last leaks its allocation. Additionally, the `DIR` destructor that releases `d_name` was gated by `#ifdef WINRT` only, so `_WIN32_WCE` builds leaked every allocation, including the last one.
## Fix
- Free the previous `dir->ent.d_name` before overwriting it in `readdir()`, matching how `~DIR()` already frees it on WinRT.
- Extend the `DIR` destructor guard from `#ifdef WINRT` to `#if defined(WINRT) || defined(_WIN32_WCE)` so the final buffer is also released under `_WIN32_WCE`.
## Checklist
- [x] I agree to contribute to the project under the Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on code under GPL or another license incompatible with OpenCV.
- [x] The PR is proposed to the proper branch (`4.x`).
- [x] There is a reference to the original bug report and related work.
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable.
- [x] The feature is well documented and sample code can be built with the project CMake.
Thanks for reviewing.
Improve fisheye extrinsic refinement for ill-conditioned views and
poorly scaled object points. Replace the Gauss-Newton update with a
Levenberg-Marquardt solver and normalize translation parameters to make
convergence less sensitive to problem scale.
Extend core performance tests - #29922
Added more cases for performance tests. It's a part of a bigger PR https://github.com/opencv/opencv/pull/29631
execution time increased by <18% per my measurements.
Duplication of https://github.com/opencv/opencv/pull/29900, but from a different fork.
`opencv_extra` PR is https://github.com/opencv/opencv_extra/pull/1410
Following was added:
| File | Test | Added |
|------|------|-------|
| `perf_arithm.cpp` | `BinaryOpTest.*` (add/subtract/multiply/absdiff/min/max/transpose2d) | types `CV_16UC1`, `CV_64FC1` |
| `perf_compare.cpp` | `compareScalar` | types `CV_16UC1`, `CV_16SC1` |
| `perf_dot.cpp` | `dot` | type `CV_64FC1` |
| `perf_flip.cpp` | `flip` (`FLIP_TYPES`) | types `CV_32FC3`, `CV_32FC4` |
| `perf_mat.cpp` | `Mat_CopyToWithMask` | types `CV_32FC3` |
| `perf_mat.cpp` | `Mat_SetToWithMask` | types `CV_8UC3, CV_8UC4, CV_16UC3, CV_16UC4, CV_32FC3, CV_32FC4` |
| `perf_norm.cpp` | `norm` | norm type `NORM_L2SQR` |
| `perf_norm.cpp` | `norm_mask` | types `CV_8UC3`, `CV_16UC3`, `CV_32FC3`; norm type `NORM_L2SQR` |
| `perf_norm.cpp` | `norm2` | norm types `NORM_L2SQR`, `NORM_RELATIVE+NORM_L2SQR` |
| `perf_norm.cpp` | `norm2_mask` | types `CV_8UC3`, `CV_16UC3`, `CV_32FC3`; norm types `NORM_L2SQR`, `NORM_RELATIVE\|NORM_L2SQR` |
| `perf_sort.cpp` | `sort`, `sorIdx` (`TYPICAL_MAT_TYPES_SORT`) | types `CV_16SC1`, `CV_32SC1`, `CV_64FC1` |
| `perf_stat.cpp` | `sum`, `mean` | types widened `{8UC1,8UC4,32FC1}` → `8U/16U/16S/32F × C1/C3/C4` |
| `perf_stat.cpp` | `mean_mask`, `meanStdDev`, `meanStdDev_mask` | types widened `{8UC1,8UC4,32FC1}` → `8U/16U/32F × C1/C3/C4` |
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [ ] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
imgproc: document integer filter2D kernels 🤖🤖🤖 - #29913
### Summary
`filter2D` accepts integer kernels, but its parameter documentation currently restricts kernels to floating-point matrices. Document the existing integer support and add parameterized accuracy coverage against `cvtest::filter2D`, the reference implementation.
The 40 cases cover five integer kernel depths (`CV_8U`, `CV_8S`, `CV_16U`, `CV_16S`, `CV_32S`), 3x3 and 13x13 kernels, and four source formats (`CV_8UC1`, `CV_32FC1`, `CV_32FC3`, `CV_64FC1`). Signed kernels include negative coefficients. Kernel depth, size, and source format are separate Google Test parameters.
Fixes#22243. Related work: #29607 was closed without merging. This follow-up accounts for its review requests about parameterized tests and avoids claiming a specific implementation path from a test name. No filtering implementation changes are included.
### Validation
Built `opencv_test_imgproc` from current 4.x (`bb9e3eff13`) on Linux with GCC 13.3.0, Release/SSE3. IPP, OpenCL, and additional CPU dispatch variants were disabled in this local build.
```
./bin/opencv_test_imgproc --gtest_filter='*Imgproc_Filter2D*:*Imgproc_FilterSupportedFormats*'
```
All 58 selected tests passed: 40 new cases and 18 existing filter accuracy/format/regression cases. `git diff --check` passed. This is targeted CPU validation, not a full imgproc suite or a GPU/backend validation claim. The tests use synthetic data and need no opencv_extra assets.
### Pull Request Readiness Checklist
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on code under GPL or another license incompatible with OpenCV.
- [x] The PR is proposed to the proper branch: 4.x for a documentation correction applicable to both maintained branches.
- [x] There is a reference to the original bug report and related work.
- [x] Accuracy tests are included; no external test data or performance change is involved.
- [x] The parameter documentation is updated and the affected test target builds.
### AI assistance
OpenAI Codex assisted with the source/review audit, patch preparation, and local validation. The commit includes an Assisted-by trailer.
imgcodecs: fix Sun Raster encoding checks 🤖🤖🤖 - #29894
SunRasterDecoder saves the header's ras_type field in m_encoding but several validation and decoding checks compared RAS_BYTE_ENCODED and RAS_FORMAT_RGB against m_type. This caused valid byte-encoded and RGB-format Sun Raster images to be rejected before their existing decoding paths could be used.
This patch uses the file encoding to make these decisions and limits byte encoding to 8-bit input and RGB-format input to the supported 24- and 32-bit paths. It also selects channel conversion from the file and requested output orders and avoids indexed-palette conversion when decoding truecolor inputs as grayscale.
The regression tests use Sun Raster files written by Netpbm's `pnmtorast -rle` and ImageMagick's SUN encoder. The matching `opencv_extra` PR (opencv/opencv_extra#1408) contains these files and lets them be inspected independently with compatible image viewers. The tests verify exact RLE literal and run values, truncated RLE input, 24- and 32-bit RGB-format input, BGR and RGB output, and grayscale conversion.
The Netpbm file also exposed an existing row-padding error in the RLE path: the decoder consumed a padding byte after even-width rows, although those rows require no padding. The patch now consumes that byte only for odd-width rows.
No public API is changed.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
Backport of #29905 to the 4.x branch, applied verbatim: after this change
`modules/python/package/cv2/__init__.py` is byte-identical to the 5.x version
(same blob 99685b3791), since both branches carried the same code here.
`import cv2` unconditionally rewrites the environment; when LD_LIBRARY_PATH
was previously unset this creates a dangling separator, i.e. an empty entry
which, per ld.so(8), resolves to the current working directory of every
subsequently spawned child process. Children may then pick up same-named
shared libraries from the CWD and fail (real-world case reported in
opencv/opencv-python#1268: a Nuitka-standalone app's `xdg-open -> kde-open`
dying with `libssl.so.3: version 'OPENSSL_3.2.0' not found`; reproduced on
4.11.0.86 - 4.14.0.94 manylinux wheels). The Windows PATH line has the same
dangling-separator pattern.
Identical to #29905: factor the prepend into _prepend_env_paths(), filter
empty entries, write the variable only when there is something to add, and
join with the old value without producing an empty entry.
The follow-up question of not mutating the process environment at all
remains tracked in #28994.
core: handle zero-sized broadcast dimensions #29878
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
<!-- Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the PR title. -->
### Changes
- Added test coverage for the zero-sized dimension case in `cv::broadcast()`.
- The test exercises the `false` branch of `_flatten_for_broadcast()`.
- Fixed division by zero when the broadcast destination has zero elements.
### Test
- `opencv_test_core.exe --gtest_filter=BroadcastTo.*`
All `BroadcastTo` tests pass.
Fixes#28910
video: fix DISOpticalFlow heap-buffer-overflow with patch_size > border_size - #29853Fixes#20185. 4.x companion to #29715.
@asmorkalov asked to retarget #29715 to 4.x since the issue reproduces there too, but that PR's branch descends from 5.x, so literally changing its base produces an unreviewable ~2M-line diff (the two branches have diverged far beyond this module). Opening a separate PR instead: same two commits, cherry-picked cleanly onto 4.x's current tip with zero conflicts (`modules/video/src/dis_flow.cpp` is byte-for-byte identical between the branches apart from this fix).
### What / why
See #29715 for the full writeup. Summary: `DISOpticalFlowImpl` pads `I1` with a fixed 16px border, while `PatchInverseSearch`'s search-position clamp lets a patch be placed up to `patch_size - 1` px outside the image -- safe only while `patch_size <= border_size`. Since `patch_size` is user-settable with no upper bound relative to the hardcoded border, `setPatchSize()` past 16 (or a large-enough temporal-candidate flow) reads past the end of the padded buffer (confirmed via AddressSanitizer).
Fix (per review on #29715): rather than growing the persistent `border_size` object field to match `patch_size` (which a reviewer correctly flagged as the wrong place for a per-call derived quantity), the search clamp itself (`i/j_lower_limit`, `i/j_upper_limit` in `PatchInverseSearch_ParBody`) now accounts for the read window needing to stay inside the *existing* padded buffer. For `patch_size <= border_size` (every built-in preset) the bounds are algebraically identical to the originals -- `border_size` itself is untouched.
### Testing
Verified this reproduces identically on 4.x: built an ASan-instrumented Debug configuration (`core+imgproc+imgcodecs+features2d+flann+video+ts`), confirmed the crash reproduces on pristine 4.x (same crash site as the original report), confirmed it's gone with this fix, and ran the new `regression_20185_patch_larger_than_border` / `regression_20185_stress` tests (5 repeated runs) plus the full `opencv_test_video` suite -- no failures attributable to this change (everything else failing needs `opencv_extra` test data not configured in this scoped build).
### PR checklist
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on code under GPL or another incompatible license.
- [x] The PR is proposed to the proper branch (4.x, per maintainer request on #29715).
- [x] Accuracy tests included (see above).
- [x] No public API/behavior change, so no documentation or sample updates needed.
pointSetBoundingRect() passed CV_32F coordinates to cvFloor() without
honouring its documented INT_MIN..INT_MAX precondition, so a point set
outside the int range collapsed to a 1x1 rect at INT_MIN, while +inf and
-inf gave INT_MIN and INT_MAX - an inverted rectangle. The extrema are
now searched for in floats, as the vectorized path already did, and
floored and saturated to the closest representable bound once, after the
reduction.
The sides are computed in int64 and clamped to INT_MAX: xmax - xmin + 1
overflows int once the extrema saturate, and did so already for a CV_32S
point set spanning the whole int range.
Fixes#29837
- Type/stride substitution only; arithmetic and evaluation order unchanged.
- RedBlackSOR: the v_extract<3>(prev, next) previous-lane construction is
replaced by an unaligned vx_load(p_next + j - 1); v_extract<3> is only
previous-lane when vlanes == 4, so this is required for wider lanes.
- HorPass keeps a strict bound, j < len - vlanes, while the other three use
j <= len - vlanes, since the vector body applies UPDATE to every lane and
the rightmost element must not receive UPDATE across the right border.
- No SVE claim: there is no SVE backend in-tree; scalable means RVV today.
- RedBlackSOR: drop the pW_next_vec load, unused since the lane-shift now
reads pW_next directly.
imgproc: Optimized Moments & perf test added - #29393
- Added dispatch for moments (CV_8U, CV_16U, CV_32F and CV_64F)
- CV_16S is kept in scalar due to regresssions observed.
- Perf test added for meanShift, CamShift and matchShapes.
- Extend Moments1 perf coverage to CV_8U alongside 16/32/64-bit depths.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
dnn: fix eltwise layer memory traversal pattern causing Windows-ARM64 slowdown - #29831
## Problem
- `EltwiseInvoker::operator()` (backs the `Eltwise` layer's SUM/PROD/DIV ops) ran slower on Windows-ARM64 than x64.
## Root cause
- The loop tiled the output into fixed-size blocks and looped over the **channel axis inside each block**.
- In NCHW layout, consecutive channels of the same sample sit far apart in memory (`planeSize * 4` bytes apart).
- So each block jumped between ~256 memory locations per tensor — roughly **768 concurrent strided streams** across the three tensors involved.
- Hardware prefetchers can only track a small, fixed number of streams and Windows-ARM64's prefetcher falls over on this pattern.
## Fix
- Walk the output buffer contiguously and track the current (sample, channel) plane from the offset and clamp each block so it never crosses into the next channel.
- This removes the inner channel loop entirely: one channel finishes before the next starts, so every tensor is read/written sequentially instead of jumping around.
- No behavior change, only the traversal order is different.
## Performance Benchmarks:
<img width="1479" height="208" alt="image" src="https://github.com/user-attachments/assets/0f878757-d988-4564-b166-96d80834c77e" />
core: vectorize masked norm/normDiff for remaining depths - #29815
### Summary
- The masked `cv::norm()` / `cv::norm(a, b)` kernels in `norm.simd.hpp` had SIMD specializations only for `uchar`, `ushort` and `float`.
- `schar`, `short`, `int`, `double` and `uchar` L2 paths uses the scalar implementation.
### Changes
- Added new vectorized implementations of MaskedNorm{Inf,L1,L2}_SIMD for schar, short, int, and double.
- Added new vectorized implementation of MaskedNormL2_SIMD<uchar, int>.
- Added new MaskedNormDiff{Inf,L1,L2}_SIMD implementations for double.
- Added cn == 4 v_load_deinterleave paths to the uchar L1/L2 kernels.
- Replaced the single f64 accumulator in MaskedNormL1_SIMD<float,double> with four independent ones, so the widening adds can overlap instead of each waiting on the previous.
### Performance Benchmarks
<img width="715" height="709" alt="image" src="https://github.com/user-attachments/assets/46afea06-df98-4dd4-a346-dc8cd4a065c9" />