Unified DNN graph fusion behind fuseChains - #30088
Requires: https://github.com/opencv/opencv_extra/pull/1417
### What This PR Does
Follow-up to #29656, which added the backend-agnostic chain mechanism. That PR introduced it and made pointwise fusion its first consumer; this one moves the rest of the fusion onto it.
- `Net::Impl::fuseBasic()` is gone. Its conv+BatchNorm, conv+activation and conv+residual branches now go through the chain mechanism.
- `fuseBN` and the six other passes now run from inside `fuseChains()`, in their existing order.
- `graph_fusion_basic.cpp` keeps `fuseBN()` and `fuseInstanceNormAffine()`.
- New `FusionOps::foldInputScale` folds a scale sitting before a conv directly into its weights, across the dense, depthwise, grouped and MLAS-prepacked layouts.
### Results
Measured on Intel Core i9-11900K
| Model | 5.x | this PR | speedup |
|-----------------------------|-----------|-----------|---------|
| inception_v2 | 15.04 ms | 11.80 ms | 1.27 |
| efficientnet-lite4 | 14.83 ms | 12.28 ms | 1.21 |
| pose_estimation_mediapipe | 5.66 ms | 5.25 ms | 1.08 |
| siglip_base_patch16_224 | 143.04 ms | 135.27 ms | 1.06 |
| densenet121 | 25.99 ms | 24.76 ms | 1.05 |
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
dnn: fold Shape of static-shape inputs in constFold() - #30065
When a graph input has a fully known shape, `constFold()` now replaces `Shape` ops on it with constants. The usual const-folding cascade then also removes the Gather/Concat/Cast/Slice ops that follow.
**What it adds**
- `graph_const_fold.cpp`: tracks known shapes and types through the graph during folding and folds `Shape` layers whose input shape is known.
- `net_impl2.cpp`: once a `Shape` has been folded, `setGraphInput()` rejects an input of a different shape instead of running with stale constants. It also fixes `tryInferGraphShapes()` to report untyped const or empty inputs as `Mat().type()`. The extra folding exposed a Resize type-check failure without this fix.
- `test_fusion.cpp`: two tests, one checking that no `Shape` layer remains after folding, one checking that a different input shape is rejected.
**Effect on graphs** (outputs identical to 5.x)
CRNN 84 → 34 layers, ViT/DeiT 361 → 274, BEiT 373 → 286, MobileViT_XS 352 → 289.
**Performance** (OCV/CPU, 32 threads; median of 5 interleaved runs per side)
| Model | 5.x, ms | PR, ms | Speedup |
|---|---|---|---|
| MobileViT_XS | 6.50 | 6.23 | 1.04x |
| BEiT_Base_Patch16_224 | 30.03 | 29.23 | 1.03x |
| VIT_Base_Patch16_224 | 27.17 | 26.53 | 1.02x |
| DeiT_Tiny_Patch16_224 | 4.90 | 4.81 | 1.02x |
| CRNN | 2.52 | 2.51 | 1.01x |
| FacePaint | 328.1 | 329.5 | 1.00x |
**Dependency**
This depends on the upcoming CUDA integration in DNN and should be merged after it.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Added Backend Agnostic fusion in DNN - #29656
# What This Adds
A way for any layer, on any backend, to absorb trailing per-element math — with the pass knowing nothing about either layer and holding no list of op types.
- **Description**: `LayerMath` (one layer's math) → `AdjacencyGraph` (hash-consed chain)
- **Contract**: two `bool` virtuals on `Layer`
- **CPU implementations**: three, behind that contract
Pointwise fusion is the first consumer, not the point.
## The Problem
```cpp
ActivationLayer* activ = dynamic_cast<ActivationLayer*>(layer_ptr);
Conv2Layer* conv = getLayer<Conv2Layer>(newprog, conv_layer_idx);
if (conv) conv->fuseActivation(layer);
```
Guest type, host type, method name all hardcoded in the pass. `Clip` never fused (fails the cast), `Gemm` never fused (no virtual), no 2-op chain fused at all, and no backend could fuse anything without a CPU-specific method.
## Results
| Model | 5.x | this PR | speedup |
|-----------------------|----------|----------|---------|
| MPHand | 2.34 ms | 1.03 ms | 2.27 |
| EfficientNet | 9.82 ms | 5.00 ms | 1.96 |
| MPPose | 5.14 ms | 2.81 ms | 1.83 |
| MobileNet_SSD_v1_ONNX | 12.58 ms | 8.38 ms | 1.50 |
| BlazeFace | 0.95 ms | 0.77 ms | 1.23 |
| DenseNet_121 | 20.56 ms | 19.08 ms | 1.08 |
| MobileNetv2_ONNX | 2.07 ms | 1.99 ms | 1.04 |
| MobileViT_XS | 6.33 ms | 6.11 ms | 1.04 |
| YuNet_320 | 1.33 ms | 1.27 ms | 1.04 |
| PPOCRv3 | 43.05 ms | 42.15 ms | 1.02 |
| BERT | 9.78 ms | 9.68 ms | 1.01 |
| BEiT_Base_Patch16_224 | 27.48 ms | 27.08 ms | 1.01 |
| DeiT_Tiny_Patch16_224 | 4.74 ms | 4.72 ms | 1.01 |
| MPPalm | 1.88 ms | 1.85 ms | 1.01 |
| SSD | 50.19 ms | 49.82 ms | 1.01 |
| YOLOv4_tiny | 7.10 ms | 7.03 ms | 1.01 |
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake