core: saturating float/double->intXY conversions in cvRound/cvFloor/cvCeil and the corresponding univ. intrinsics. - #30007
Merge pull request #30007 from vpisarev:fix_flt2int
Fixes#28557 (supersedes the tail-only fix from #29895).
### The problem
On x86 `cvtss2si`/`cvtsd2si`/`cvtps2dq` return the "integer indefinite" value `0x80000000` (`INT_MIN`) for any out-of-range input, so
```cpp
cvRound(3e9); // INT_MIN
saturate_cast<ushort>(60000.f*60000.f); // 0 instead of 65535 (#28557: the scalar tail of cv::multiply)
v_round(v_float32(1e10f)); // INT_MIN lanes
```
The problem is not x86-only. `cvRound()` on aarch64, riscv64 and loongarch64 went through `(int)lrint()`, which the compilers implement as a 64-bit conversion followed by truncation to 32 bits, i.e. large values *wrap* instead of saturating (`fcvtzs x0` + `mov w0`, `fcvt.l.d` + `sext.w`). NEON `v_round/v_floor/v_ceil/v_trunc(v_float64x2)` narrowed with a wrapping `vmovn`, and `saturate_cast<unsigned/int64/uint64>(float/double)` were plain UB above the target range.
### The fix
**`fast_math.hpp`**
- `cvRound()`, `cvFloor()`, `cvCeil()` saturate on every platform. New `cvTrunc()` (round towards zero, saturating; the scalar counterpart of `v_trunc`) and `cvRound64()` (round-half-to-even to `int64`, saturating).
- x86 (SSE2): the input is clamped from above (`min`) before `cvt*`; the "indefinite" value is already the correct result below `INT_MIN`. `cvFloor` also clamps from below because of the "-1" correction.
- aarch64 (`__GNUC__`/clang): one-instruction inline asm `fcvtns/fcvtms/fcvtps/fcvtzs` (GCC's `arm_neon.h` has no scalar f64→s32 ACLE functions, and inline asm needs no header).
- riscv64: `fcvt.w.{s,d}` / `fcvt.l.{s,d}` with an explicit rounding mode (`rne`, `rdn`, `rup`, `rtz`).
- loongarch64: `ftint*.w.{s,d}` + `movfr2gr.s` (the old `.l.d` + `movfr2gr.d` wrapped); `cvRound` got a branch too.
- MSVC ARM64: the 64-bit results are clamped before narrowing.
- Portable `#else` branch: hand-written clamp before the conversion, no new headers.
- Exact semantics: `double` saturates to `INT_MIN`/`INT_MAX` exactly. `INT_MAX` is not representable as `float`, so `float` inputs `>= 2^31` give **2147483520** (the largest float below 2^31, the clamp value) where the instruction does not saturate by itself (x86, portable) and `INT_MAX` where it does (ARM, RISC-V, LoongArch). Both are accepted by the tests and documented.
- **NaN handling is out of scope in this PR**.
**`saturate.hpp`**
- `float/double → unsigned/int64/uint64` now saturate and use round-half-to-even via `cvRound64()` (no libm `round()` call; OpenCV builds without `-ffast-math`, where `rint()`/`round()` are real calls).
- `int/int64 → schar/short/int` range checks are done in unsigned arithmetic. The old `(unsigned)(v - SHRT_MIN)` overflowed (UB) for `v > INT_MAX - 32768`; such values were unreachable before but are returned by the saturating `cvRound()` now, and GCC 15 actually exploits the UB (derives `v <= INT_MAX-32768` and folds neighbouring comparisons).
**Universal intrinsics**
- SSE, AVX2, AVX-512: clamp before `cvt` in `v_round/v_floor/v_ceil/v_trunc` (f32 and f64, `v_round(a, b)` included). AVX2 keeps `_mm256_floor_ps/_ceil_ps`.
- NEON f64: saturating narrow (`vqmovn_s64`); `v_floor/v_ceil` use `fcvtms/fcvtps` directly; `v_trunc` used `vcvtaq` (round-to-nearest-away) instead of truncation.
- WASM f32: clamp before the `-1/+1` corrections (they wrapped `INT_MIN`); `v_round` was `trunc_sat(a + 0.5)` (wrong for negatives and ties), now `f32x4.nearest`; f64 `v_trunc` via `cvTrunc`.
- MSA f64: clamp before `pckev.w`, which takes the low 32 bits of the saturated int64.
- RVV f16 `v_floor`: rounding mode `2` (RDN) instead of `1` (RTZ).
- VSX, LSX/LASX, RVV f32/f64, RVV 0.7.1: untouched, the conversion instructions saturate by ISA definition.
- `intrin_cpp.hpp`: `v_trunc` via `cvTrunc`; docs mention the saturation.
`arithm.simd.hpp` is not modified: `Core_Arithm.mul_overflow_28557` passes because the scalar tail's `saturate_cast<ushort>(float)` saturates now.
### Tests
- `Core_Arithm.mul_overflow_28557` re-enabled.
- `Core_Arithm.mul_overflow_tail_and_inplace`: 8U/8S/16U/16S, row lengths 1..70 (vector body + scalar tail), out-of-place and in-place `multiply`, `mul`, `pow(x, 2)`.
- `Core_ConvertTo.float_overflow_saturation`: 32F/64F → 8U/8S/16U/16S/32S/32U/64S/64U with `±1e10`, `±1e30`, `±inf`, `2147483647.5`, `x.5` etc.; vector body vs. scalar reference, extreme inputs must hit the type limits.
- `Core_FastMath.SaturatingRoundingOps`: boundary table (`2147483647.5`, `-2147483648.5`, `2147483520f`, `-2147483904f`, `±DBL_MAX`, `±inf`, ...) plus a 200k-value log-uniform random sweep against `nearbyint/floor/ceil/trunc` clamped to `int`.
- `Core_FastMath.Round64`: boundaries around `±2^63`, `.5` cases, random sweep.
- `Core_SaturateCast.FloatToIntSaturation`, `RoundHalfToEven`, `IntToNarrowerIntBoundaries` (regression for the signed-overflow UB).
- `test_intrin_utils.hpp` (`test_float_math`, `test_round_pair_f64`): overflow/boundary lanes for f32, f64 and f16, compared against the scalar functions and checked against explicit `INT_MIN` / `[2147483520, INT_MAX]` / `INT_MAX` limits. Runs on every backend and dispatch level.
### Verification
- x86-64, GCC 15.2, `CPU_BASELINE=SSE4_2 CPU_DISPATCH=AVX2,AVX512_SKX`: full `opencv_test_core` passes (16864 tests), i.e. SSE4.2 baseline, AVX2 dispatch and the CPP emulator paths. AVX-512 is compile-checked only (no AVX-512 hardware here).
- The standalone scalar checks also pass with the portable `#else` branch forced (`-U__SSE2__`) and under `-fsanitize=undefined`.
- aarch64 (`aarch64-linux-gnu-gcc-14`) and riscv64 (`riscv64-linux-gnu-gcc-14 -march=rv64gcv_zvfh`): compile-checked, generated code inspected (`fcvtns w0, d0`, `fcvt.w.d a0, fa0, rne`, `sqxtn`, `vfcvt.x.f.v`, ...). Not run (no hardware/qemu).
- **Not verified at all** (please watch CI): LoongArch (`ftint*.w.d`/`ftintrne.*` asm), WASM (`wasm_f32x4_nearest`), MSA, MSVC ARM64 (`vcvtd_s64_f64`, `vcvts_s32_f32`, `vcvtnd_s64_f64`).
### Behaviour changes worth noting
- `saturate_cast<unsigned/int64/uint64>(x.5)` now rounds half to even (was half away from zero), consistent with all the other integer targets.
- `cvRound(float)` for inputs `>= 2^31` returns **2147483520** on x86 (was `INT_MIN`); `INT_MAX` on ARM/RISC-V. It would be noticeably slower to implement true saturation to `INT_MAX`. Note that around `INT_MAX` float's cannot represent the integer's exactly anyway.
- `saturate_cast<int>(1e10)` is `INT_MAX` (documentation used to say "no clipping is done for 32-bit integers").
### Pull Request Readiness Checklist
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on code under GPL or another license incompatible with OpenCV.
- [x] The PR is proposed to the proper branch (`5.x`).
- [x] There is a reference to the original bug report and related work: #28557, #29895.
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable: accuracy tests added, no test data needed.
- [x] The feature is well documented and sample code can be built with the project CMake: doxygen comments of `cvRound`/`cvFloor`/`cvCeil`/`cvTrunc`/`cvRound64`, `saturate_cast` and the intrinsics conversion group updated.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01KVnDjvrFtxcyw7RFzSPews
Remove deprecated CommaInitializer API - #29948
This goes hand in hand with https://github.com/opencv/opencv_contrib/pull/4217
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
Better Durand-Kerner Initialization #29109
While investigating issue #23644, I have found [this paper](https://link.springer.com/article/10.1007/BF01935059) which presents a good initialization for the Durand-Kerner algorithm. Basically the idea is to put the initial points equidistantly on a circle on the complex plane. The radius of the circle is computed as
<img width="607" height="178" alt="image" src="https://github.com/user-attachments/assets/ea31b002-c924-4b93-9334-3e59597c896b" />
Note that the $a_i$ coefficients in that paper are reversed compared to OpenCV. That's where the `(n - i)` in the code comes from.
I have implemented just the mean of the $u_i$'s for the sake of simplicity. That's already enough to make the algorithm converge in all cases I have tested. I have used this to test for convergence for many polynomials of order 2 and 4 and coefficients of different magnitudes:
```cpp
TEST(Core_SolvePoly, large_test)
{
cv::Mat_<float> coefs3(1,3);
cv::Mat_<float> coefs5(1,5);
cv::Mat r;
double prec;
for (int c0 = -20; c0 <= 20; c0++)
{
coefs3.at<float>(0) = c0;
for (int c1 = -20; c1 <= 20; c1++)
{
coefs3.at<float>(1) = c1;
for (int c2 = -20; c2 <= 20; c2++)
{
coefs3.at<float>(2) = c2;
prec = cv::solvePoly(coefs3, r);
EXPECT_LE(prec, 1e-6);
}
}
}
for (int c0 = -10; c0 <= 10; c0++)
{
coefs5.at<float>(0) = c0;
for (int c1 = -10; c1 <= 10; c1++)
{
coefs5.at<float>(1) = c1;
for (int c2 = -10; c2 <= 10; c2++)
{
coefs5.at<float>(2) = c2;
for (int c3 = -10; c3 <= 10; c3++)
{
coefs5.at<float>(3) = c3;
for (int c4 = -10; c4 <= 10; c4++)
{
coefs5.at<float>(4) = c4;
prec = cv::solvePoly(coefs5, r);
EXPECT_LE(prec, 1e-2);
}
}
}
}
}
for (int i = -10; i < 10; i++)
{
coefs3.at<float>(0) = pow(2, i);
for (int j = -10; j < 10; j++)
{
coefs3.at<float>(1) = pow(2, j);
for (int k = -10; k < 10; k++)
{
coefs3.at<float>(2) = pow(2, k);
prec = cv::solvePoly(coefs3, r);
EXPECT_LE(prec, 1e-6);
}
}
}
}
```
This test passes, but I have not committed it because it runs for a couple of seconds.
This fixes#23644 and replaces #29055. I have checked #29055 and it does not pass the test above. It seems to be optimized to the precise polynomial of #23644.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
core: fix solveCubic numerical instability via coefficient normalization (fixes#27748) #28117
Summary
This PR fixes numerical instability in `cv::solveCubic` when the leading coefficient `a` is non-zero but extremely small relative to other coefficients (Issue #27748).
It introduces a **normalization step** that scales all coefficients by their maximum magnitude before solving. This ensures robust detection of when the equation should degenerate to a quadratic solver, without breaking valid cubic equations that happen to have small coefficients (e.g., scaled by 1e-9).
The Problem (Issue #27748)
The previous implementation checked `if (a == 0)` to decide whether to use the cubic or quadratic formula.
- When `a` is extremely small (e.g., 1e-17) but not exactly zero, and other coefficients are normal (e.g., 5.0), the standard cubic formula suffers from catastrophic cancellation and overflow, producing incorrect roots (e.g., 1e14).
The Fix
1. Normalization: The solver now finds `max_coeff = max(|a|, |b|, |c|, |d|)` and scales all coefficients by `1.0 / max_coeff`.
2. Relative Threshold: It then checks `if (abs(a) < epsilon)` on the *normalized* coefficients.
Why this is better than previous attempts
In a previous attempt (PR #28057), a simple absolute check `abs(a) < epsilon` was proposed. That approach was rejected because it failed for scaled equations.
Fixes#27748
Improve solveCubic accuracy #27347
### Pull Request Readiness Checklist
Fix#27323
```
2e-13 * x^3 + x^2 - 2 * x + 1 = 0 -> x^3 + 5e12 * x^2 - 1e13 * x + 5e12 = 0
```
The problem that coefficients have quite big magnitudes and current calculations are subject to round-off error
```
Q = (a1 * a1 - 3 * a2) * (1./9)
R = (2 * a1 * a1 * a1 - 9 * a1 * a2 + 27 * a3) * (1./54)
Qcubed = Q * Q * Q = a1^6/729 - (a1^4 a2)/81 + (a1^2 a2^2)/27 - a2^3/27
R * R = R^2 = a1^6/729 - (a1^4 a2)/81 + (a1^2 a2^2)/36 + (a1^3 a3)/27 - (a1 a2 a3)/6 + a3^2/4
d = Qcubed - R * R
```
Let `a1`, `a2`, `a3` have quite big same magnitudes, then we see that `Qcubed` and `R * R` have same terms `a1^6/729` and `-(a1^4 a2)/81` (which will be reduced in `d`), but they level out the other terms (these terms have `6`th and `5`th degree and other terms - less or equal than `4`th degree).
So, if these terms will participate in the calculation, this will lead to a huge round-off error.
But if we expand the expression, then round-off error should be less
```
d = Qcubed - R * R = 1/108 (a1^2 a2^2 - 4 a2^3 - 4 a1^3 a3 + 18 a1 a2 a3 - 27 a3^2)
```
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Add tests for solveCubic #27331
### Pull Request Readiness Checklist
Related to #27323
I found only randomized tests with number of roots always equal to `1` or `3`, `x^3 = 0` and some simple test for Java and Swift.
Obviously, they don't cover all cases (implementation has strong branching and number of roots can be equal to `-1`, `0` and `2` additionally).
So, I think it will be useful to try explicitly cover more cases (and implementation branches correspondingly)
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
2x more accurate float => bfloat conversion #26321
There is a magic trick to make float => bfloat conversion more accurate (_original reference needed, is it done this way in PyTorch?_). In simplified form it looks like:
```
uint16_t f2bf(float x) {
union {
unsigned u;
float f;
} u;
u.f = x;
// return (uint16_t)(u.u >> 16); <== the old method before this patch
return (uint16_t)((u.u + 0x8000) >> 16);
}
```
it works correctly for almost all valid floating-point values, positive, zero or negative, and even for some extreme cases, like `+/-inf`, `nan` etc. The addition of `0x8000` to integer representation of 32-bit float before retrieving the highest 16 bits reduces the rounding error by ~2x.
The slight problem with this improved method is that the numbers very close to or equal to `+/-FLT_MAX` are mistakenly converted to `+/-inf`, respectively.
This patch implements improved algorithm for `float => bfloat` conversion in scalar and vector form; it fixes the above-mentioned problem using some extra bit magic, i.e. 0x8000 is not added to very big (by absolute value) numbers:
```
// the actual implementation is more efficient,
// without conditions or floating-point operations, see the source code
return (uint16_t)(u.u + (fabsf(x) <= big_threshold ? 0x8000 : 0)) >> 16);
```
The corresponding test has been added as well and this is output from the test:
```
[----------] 1 test from Core_BFloat
[ RUN ] Core_BFloat.convert
maxerr0 = 0.00774842, mean0 = 0.00190643, stddev0 = 0.00186063
maxerr1 = 0.00389057, mean1 = 0.000952614, stddev1 = 0.000931268
[ OK ] Core_BFloat.convert (7 ms)
```
Here `maxerr0, mean0, stddev0` are for the original method and `maxerr1, mean1, stddev1` are for the new method. As you can see, there is a significant improvement in accuracy.
**Note:**
_Actually, on ~32,000,000 random FP32 numbers with uniformly distributed sign, exponent and mantissa the new method is always at least as accurate as the old one._
The test also checks all the corner cases, where we see no degradation either vs the original method.
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
C-API cleanup: apps, imgproc_c and some constants #25075
Merge with https://github.com/opencv/opencv_contrib/pull/3642
* Removed obsolete apps - traincascade and createsamples (please use older OpenCV versions if you need them). These apps relied heavily on C-API
* removed all mentions of imgproc C-API headers (imgproc_c.h, types_c.h) - they were empty, included core C-API headers
* replaced usage of several C constants with C++ ones (error codes, norm modes, RNG modes, PCA modes, ...) - most part of this PR (split into two parts - all modules and calib+3d - for easier backporting)
* removed imgproc C-API headers (as separate commit, so that other changes could be backported to 4.x)
Most of these changes can be backported to 4.x.
Added clapack
* bring a small subset of Lapack, automatically converted to C, into OpenCV
* added missing lsame_ prototype
* * small fix in make_clapack script
* trying to fix remaining CI problems
* fixed character arrays' initializers
* get rid of F2C_STR_MAX
* * added back single-precision versions for QR, LU and Cholesky decompositions. It adds very little extra overhead.
* added stub version of sdesdd.
* uncommented calls to all the single-precision Lapack functions from opencv/core/src/hal_internal.cpp.
* fixed warning from Visual Studio + cleaned f2c runtime a bit
* * regenerated Lapack w/o forward declarations of intrinsic functions (such as sqrt(), r_cnjg() etc.)
* at once, trailing whitespaces are removed from the generated sources, just in case
* since there is no declarations of intrinsic functions anymore, we could turn some of them into inline functions
* trying to eliminate the crash on ARM
* fixed API and semantics of s_copy
* * CLapack has been tested successfully. It's now time to restore the standard LAPACK detection procedure
* removed some more trailing whitespaces
* * retained only the essential stuff in CLapack
* added checks to lapack calls to gracefully return "not implemented" instead of returning invalid results with "ok" status
* disabled warning when building lapack
* cmake: update LAPACK detection
Co-authored-by: Alexander Alekhin <alexander.a.alekhin@gmail.com>
Add a basic sanity test to verify the rounding functions
work as expected.
Likewise, extend the rounding performance test to cover the
additional float -> int fast math functions.
* rewrote Mat::convertTo() and convertScaleAbs() to wide universal intrinsics; added always-available and SIMD-optimized FP16<=>FP32 conversion
* fixed compile warnings
* fix some more compile errors
* slightly relaxed accuracy threshold for int->float conversion (since we now do it using single-precision arithmetics, not double-precision)
* fixed compile errors on iOS, Android and in the baseline C++ version (intrin_cpp.hpp)
* trying to fix ARM-neon builds
* trying to fix ARM-neon builds
* trying to fix ARM-neon builds
* trying to fix ARM-neon builds