core: saturating float/double->intXY conversions in cvRound/cvFloor/cvCeil and the corresponding univ. intrinsics. - #30007
Merge pull request #30007 from vpisarev:fix_flt2int
Fixes#28557 (supersedes the tail-only fix from #29895).
### The problem
On x86 `cvtss2si`/`cvtsd2si`/`cvtps2dq` return the "integer indefinite" value `0x80000000` (`INT_MIN`) for any out-of-range input, so
```cpp
cvRound(3e9); // INT_MIN
saturate_cast<ushort>(60000.f*60000.f); // 0 instead of 65535 (#28557: the scalar tail of cv::multiply)
v_round(v_float32(1e10f)); // INT_MIN lanes
```
The problem is not x86-only. `cvRound()` on aarch64, riscv64 and loongarch64 went through `(int)lrint()`, which the compilers implement as a 64-bit conversion followed by truncation to 32 bits, i.e. large values *wrap* instead of saturating (`fcvtzs x0` + `mov w0`, `fcvt.l.d` + `sext.w`). NEON `v_round/v_floor/v_ceil/v_trunc(v_float64x2)` narrowed with a wrapping `vmovn`, and `saturate_cast<unsigned/int64/uint64>(float/double)` were plain UB above the target range.
### The fix
**`fast_math.hpp`**
- `cvRound()`, `cvFloor()`, `cvCeil()` saturate on every platform. New `cvTrunc()` (round towards zero, saturating; the scalar counterpart of `v_trunc`) and `cvRound64()` (round-half-to-even to `int64`, saturating).
- x86 (SSE2): the input is clamped from above (`min`) before `cvt*`; the "indefinite" value is already the correct result below `INT_MIN`. `cvFloor` also clamps from below because of the "-1" correction.
- aarch64 (`__GNUC__`/clang): one-instruction inline asm `fcvtns/fcvtms/fcvtps/fcvtzs` (GCC's `arm_neon.h` has no scalar f64→s32 ACLE functions, and inline asm needs no header).
- riscv64: `fcvt.w.{s,d}` / `fcvt.l.{s,d}` with an explicit rounding mode (`rne`, `rdn`, `rup`, `rtz`).
- loongarch64: `ftint*.w.{s,d}` + `movfr2gr.s` (the old `.l.d` + `movfr2gr.d` wrapped); `cvRound` got a branch too.
- MSVC ARM64: the 64-bit results are clamped before narrowing.
- Portable `#else` branch: hand-written clamp before the conversion, no new headers.
- Exact semantics: `double` saturates to `INT_MIN`/`INT_MAX` exactly. `INT_MAX` is not representable as `float`, so `float` inputs `>= 2^31` give **2147483520** (the largest float below 2^31, the clamp value) where the instruction does not saturate by itself (x86, portable) and `INT_MAX` where it does (ARM, RISC-V, LoongArch). Both are accepted by the tests and documented.
- **NaN handling is out of scope in this PR**.
**`saturate.hpp`**
- `float/double → unsigned/int64/uint64` now saturate and use round-half-to-even via `cvRound64()` (no libm `round()` call; OpenCV builds without `-ffast-math`, where `rint()`/`round()` are real calls).
- `int/int64 → schar/short/int` range checks are done in unsigned arithmetic. The old `(unsigned)(v - SHRT_MIN)` overflowed (UB) for `v > INT_MAX - 32768`; such values were unreachable before but are returned by the saturating `cvRound()` now, and GCC 15 actually exploits the UB (derives `v <= INT_MAX-32768` and folds neighbouring comparisons).
**Universal intrinsics**
- SSE, AVX2, AVX-512: clamp before `cvt` in `v_round/v_floor/v_ceil/v_trunc` (f32 and f64, `v_round(a, b)` included). AVX2 keeps `_mm256_floor_ps/_ceil_ps`.
- NEON f64: saturating narrow (`vqmovn_s64`); `v_floor/v_ceil` use `fcvtms/fcvtps` directly; `v_trunc` used `vcvtaq` (round-to-nearest-away) instead of truncation.
- WASM f32: clamp before the `-1/+1` corrections (they wrapped `INT_MIN`); `v_round` was `trunc_sat(a + 0.5)` (wrong for negatives and ties), now `f32x4.nearest`; f64 `v_trunc` via `cvTrunc`.
- MSA f64: clamp before `pckev.w`, which takes the low 32 bits of the saturated int64.
- RVV f16 `v_floor`: rounding mode `2` (RDN) instead of `1` (RTZ).
- VSX, LSX/LASX, RVV f32/f64, RVV 0.7.1: untouched, the conversion instructions saturate by ISA definition.
- `intrin_cpp.hpp`: `v_trunc` via `cvTrunc`; docs mention the saturation.
`arithm.simd.hpp` is not modified: `Core_Arithm.mul_overflow_28557` passes because the scalar tail's `saturate_cast<ushort>(float)` saturates now.
### Tests
- `Core_Arithm.mul_overflow_28557` re-enabled.
- `Core_Arithm.mul_overflow_tail_and_inplace`: 8U/8S/16U/16S, row lengths 1..70 (vector body + scalar tail), out-of-place and in-place `multiply`, `mul`, `pow(x, 2)`.
- `Core_ConvertTo.float_overflow_saturation`: 32F/64F → 8U/8S/16U/16S/32S/32U/64S/64U with `±1e10`, `±1e30`, `±inf`, `2147483647.5`, `x.5` etc.; vector body vs. scalar reference, extreme inputs must hit the type limits.
- `Core_FastMath.SaturatingRoundingOps`: boundary table (`2147483647.5`, `-2147483648.5`, `2147483520f`, `-2147483904f`, `±DBL_MAX`, `±inf`, ...) plus a 200k-value log-uniform random sweep against `nearbyint/floor/ceil/trunc` clamped to `int`.
- `Core_FastMath.Round64`: boundaries around `±2^63`, `.5` cases, random sweep.
- `Core_SaturateCast.FloatToIntSaturation`, `RoundHalfToEven`, `IntToNarrowerIntBoundaries` (regression for the signed-overflow UB).
- `test_intrin_utils.hpp` (`test_float_math`, `test_round_pair_f64`): overflow/boundary lanes for f32, f64 and f16, compared against the scalar functions and checked against explicit `INT_MIN` / `[2147483520, INT_MAX]` / `INT_MAX` limits. Runs on every backend and dispatch level.
### Verification
- x86-64, GCC 15.2, `CPU_BASELINE=SSE4_2 CPU_DISPATCH=AVX2,AVX512_SKX`: full `opencv_test_core` passes (16864 tests), i.e. SSE4.2 baseline, AVX2 dispatch and the CPP emulator paths. AVX-512 is compile-checked only (no AVX-512 hardware here).
- The standalone scalar checks also pass with the portable `#else` branch forced (`-U__SSE2__`) and under `-fsanitize=undefined`.
- aarch64 (`aarch64-linux-gnu-gcc-14`) and riscv64 (`riscv64-linux-gnu-gcc-14 -march=rv64gcv_zvfh`): compile-checked, generated code inspected (`fcvtns w0, d0`, `fcvt.w.d a0, fa0, rne`, `sqxtn`, `vfcvt.x.f.v`, ...). Not run (no hardware/qemu).
- **Not verified at all** (please watch CI): LoongArch (`ftint*.w.d`/`ftintrne.*` asm), WASM (`wasm_f32x4_nearest`), MSA, MSVC ARM64 (`vcvtd_s64_f64`, `vcvts_s32_f32`, `vcvtnd_s64_f64`).
### Behaviour changes worth noting
- `saturate_cast<unsigned/int64/uint64>(x.5)` now rounds half to even (was half away from zero), consistent with all the other integer targets.
- `cvRound(float)` for inputs `>= 2^31` returns **2147483520** on x86 (was `INT_MIN`); `INT_MAX` on ARM/RISC-V. It would be noticeably slower to implement true saturation to `INT_MAX`. Note that around `INT_MAX` float's cannot represent the integer's exactly anyway.
- `saturate_cast<int>(1e10)` is `INT_MAX` (documentation used to say "no clipping is done for 32-bit integers").
### Pull Request Readiness Checklist
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on code under GPL or another license incompatible with OpenCV.
- [x] The PR is proposed to the proper branch (`5.x`).
- [x] There is a reference to the original bug report and related work: #28557, #29895.
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable: accuracy tests added, no test data needed.
- [x] The feature is well documented and sample code can be built with the project CMake: doxygen comments of `cvRound`/`cvFloor`/`cvCeil`/`cvTrunc`/`cvRound64`, `saturate_cast` and the intrinsics conversion group updated.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01KVnDjvrFtxcyw7RFzSPews
Remove deprecated CommaInitializer API - #29948
This goes hand in hand with https://github.com/opencv/opencv_contrib/pull/4217
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
core: fix overflow in borderInterpolate BORDER_WRAP path - #29884
### Description
Fixes an integer overflow/underflow bug in `cv::borderInterpolate()` when using `BORDER_WRAP`. For extreme values of `p` (e.g. `INT_MIN` or `INT_MAX`), the previous implementation performed the modulo operation directly on `int`, which can invoke undefined behavior on overflow and produce a result outside the valid `[0, len)` range.
The fix widens the intermediate calculation to `int64_t` before taking the modulo, then adjusts for negative results and narrows back to `int` only once the value is confirmed to be in range.
### Changes
- `modules/core/src/copy.cpp`: use 64-bit intermediate arithmetic in the `BORDER_WRAP` branch of `borderInterpolate()` to avoid overflow.
- `modules/core/test/test_misc.cpp`: add `Core_BorderInterpolate.wrap_no_overflow_29232`, a regression test that exercises `BORDER_WRAP` with `INT_MIN` and `INT_MAX` and asserts the result stays within `[0, len)`.
Fixes#29232
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test for the patch (added in test_misc.cpp); no performance test needed as this is a bug fix with negligible performance impact.
- [x] The feature is well documented and sample code can be built with the project CMake
Imgproc: use double to determine whether the corners points are within src #26022close#26016
Related https://github.com/opencv/opencv_contrib/pull/3778
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Removed all pre-C++11 code, workarounds, and branches #23736
This removes a bunch of pre-C++11 workrarounds that are no longer necessary as C++11 is now required.
It is a nice clean up and simplification.
* No longer unconditionally #include <array> in cvdef.h, include explicitly where needed
* Removed deprecated CV_NODISCARD, already unused in the codebase
* Removed some pre-C++11 workarounds, and simplified some backwards compat defines
* Removed CV_CXX_STD_ARRAY
* Removed CV_CXX_MOVE_SEMANTICS and CV_CXX_MOVE
* Removed all tests of CV_CXX11, now assume it's always true. This allowed removing a lot of dead code.
* Updated some documentation consequently.
* Removed all tests of CV_CXX11, now assume it's always true
* Fixed links.
---------
Co-authored-by: Maksim Shabunin <maksim.shabunin@gmail.com>
Co-authored-by: Alexander Smorkalov <alexander.smorkalov@xperience.ai>
Skip test on SkipTestException at fixture's constructor (version 2) #24250
### Pull Request Readiness Checklist
Another version of https://github.com/opencv/opencv/pull/24186 (reverted by https://github.com/opencv/opencv/pull/24223). Current implementation cannot handle skip exception at `static void SetUpTestCase` but works on `virtual void SetUp`.
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Skip test on SkipTestException at fixture's constructor
* Skip test on SkipTestException at fixture's constructor
* Add warning supression
* Skip Python tests if no test file found
* Skip instances of test fixture with exception at SetUpTestCase
* Skip test with exception at SetUp method
* Try remove warning disable
* Add CV_NORETURN
* Remove FAIL assertion
* Use findDataFile to throw Skip exception
* Throw exception conditionally
* started working on adding 32u, 64u, 64s, bool and 16bf types to OpenCV
* core & imgproc tests seem to pass
* fixed a few compile errors and test failures on macOS x86
* hopefully fixed some compile problems and test failures
* fixed some more warnings and test failures
* trying to fix small deviations in perf_core & perf_imgproc by revering randf_64f to exact version used before
* trying to fix behavior of the new OpenCV with old plugins; there is (quite strong) assumption that video capture would give us frames with depth == CV_8U (0) or CV_16U (2). If depth is > 7 then it means that the plugin is built with the old OpenCV. It needs to be recompiled, of course and then this hack can be removed.
* try to repair the case when target arch does not have FP64 SIMD
* 1. fixed bug in itoa() found by alalek
2. restored ==, !=, > and < univ. intrinsics on ARM32/ARM64.
Fow now, it is possible to define valid rectangle for which some
functions overflow (e.g. br(), ares() ...).
This patch fixes the intersection operator so that it works with
any rectangle.
* implements https://github.com/opencv/opencv/issues/19147
* CAUTION: this PR will only functions safely in the
4+ branches that already include PR 19029
* CAUTION: this PR requires thread-safe startup of the alloc.cpp
translation unit as implemented in PR 19029
- removed tr1 usage (dropped in C++17)
- moved includes of vector/map/iostream/limits into ts.hpp
- require opencv_test + anonymous namespace (added compile check)
- fixed norm() usage (must be from cvtest::norm for checks) and other conflict functions
- added missing license headers