Files
opencv/samples
Jaivardhan Bhola 193f8c4c5d Merge pull request #29618 from jaivardhan-bhola:vlm-ocr-support-engine-new
Added support for Granite Docling 258M and PaddleOCR-VL - #29618

Adds `com.microsoft::GroupQueryAttention` to the engine-new ONNX importer, required by GraniteDocling-258M's ,end-to-end samples for GraniteDocling-258M and PaddleOCR-VL-1.5. Verified on Linux x86_64 / GCC / CPU

**companion opencv_extra PR** [Vlm ocr support engine new opencv/opencv_extra#1412](https://github.com/opencv/opencv_extra/pull/1412) for test data


upstream/5.x already lowers `GroupQueryAttention` to `AttentionOnnxAi`, but that path rejects `do_rotary=1`:

```cpp
CV_CheckEQ(params.get<int>("do_rotary", 0), 0, "GroupQueryAttention: do_rotary=1 is not supported");
```

GraniteDocling-258M's export sets it. Confirmed two ways: building a Llama checkpoint from its published `text_config` (9 heads / 3 KV heads, `rope_theta=100000`) and running onnxruntime-genai's `model_builder`; and reading the community export at `onnx-community/granite-docling-258M-ONNX`, which carries 30 `GroupQueryAttention` nodes with `do_rotary=1`, `num_heads=9`, `kv_num_heads=3`, `head_dim=64`.


## Changes

**Importer** (`onnx_importer2.cpp`, +49)

Routing as above, plus the input validation the op needs: packed QKV rejected, past_key / past_value required as a pair, `do_rotary=1` requires cos/sin, unsupported trailing inputs rejected, and `seqlens_k` shorter than the buffer recognised as a shared-buffer export rather than silently mis-attended.

**Layer** (`group_query_attention_layer.cpp`, +380)

Attention is built on `fastGemmBatch` and `fused_softmax_softcap_mask`, the same kernels `attention_onnxai_layer.cpp` uses. Grouping falls out of the batched-GEMM offset table (query head → `h / groupSize`), so no broadcast copy is needed, and the second GEMM writes straight into the `[B, S, num_heads*D]` output. Supports rotary (full and partial), sliding window, softcap, growing and preallocated caches, and optional `present_*` outputs.

**Core** (`persistence_json.cpp`, +92)

`\uXXXX` escapes are UTF-16 code units, so anything above the BMP arrives as a surrogate pair. These were encoded as two separate 3-byte sequences CESU-8, which is not valid UTF-8 - for keys and values alike. Now recombined into one 4-byte sequence, with unpaired surrogates rejected.


**Tests / perf / samples**

`test_layer_group_query_attention.cpp` (+93), `test_tokenizer.cpp` (+28), `test_tokenizer.py` (+42), `perf_layer.cpp` (+125: `Layer_GroupQueryAttention` MHA_ShortCache and Grouped_WithCache, plus `Layer_Sign`), and two samples (`granite_docling_inference.py`, `paddleocr_vl_inference.py`).


### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
2026-10-01 18:03:55 +03:00
..
2025-04-02 13:45:08 -07:00
2021-12-30 21:43:45 +00:00
2026-06-22 20:38:16 -04:00
2026-09-23 14:28:55 +02:00
2025-07-22 09:47:19 +03:00