Commit Graph
3 Commits
Author SHA1 Message Date
Jaivardhan Bhola 193f8c4c5d Merge pull request #29618 from jaivardhan-bhola:vlm-ocr-support-engine-new
Added support for Granite Docling 258M and PaddleOCR-VL - #29618

Adds `com.microsoft::GroupQueryAttention` to the engine-new ONNX importer, required by GraniteDocling-258M's ,end-to-end samples for GraniteDocling-258M and PaddleOCR-VL-1.5. Verified on Linux x86_64 / GCC / CPU

**companion opencv_extra PR** [Vlm ocr support engine new opencv/opencv_extra#1412](https://github.com/opencv/opencv_extra/pull/1412) for test data


upstream/5.x already lowers `GroupQueryAttention` to `AttentionOnnxAi`, but that path rejects `do_rotary=1`:

```cpp
CV_CheckEQ(params.get<int>("do_rotary", 0), 0, "GroupQueryAttention: do_rotary=1 is not supported");
```

GraniteDocling-258M's export sets it. Confirmed two ways: building a Llama checkpoint from its published `text_config` (9 heads / 3 KV heads, `rope_theta=100000`) and running onnxruntime-genai's `model_builder`; and reading the community export at `onnx-community/granite-docling-258M-ONNX`, which carries 30 `GroupQueryAttention` nodes with `do_rotary=1`, `num_heads=9`, `kv_num_heads=3`, `head_dim=64`.


## Changes

**Importer** (`onnx_importer2.cpp`, +49)

Routing as above, plus the input validation the op needs: packed QKV rejected, past_key / past_value required as a pair, `do_rotary=1` requires cos/sin, unsupported trailing inputs rejected, and `seqlens_k` shorter than the buffer recognised as a shared-buffer export rather than silently mis-attended.

**Layer** (`group_query_attention_layer.cpp`, +380)

Attention is built on `fastGemmBatch` and `fused_softmax_softcap_mask`, the same kernels `attention_onnxai_layer.cpp` uses. Grouping falls out of the batched-GEMM offset table (query head → `h / groupSize`), so no broadcast copy is needed, and the second GEMM writes straight into the `[B, S, num_heads*D]` output. Supports rotary (full and partial), sliding window, softcap, growing and preallocated caches, and optional `present_*` outputs.

**Core** (`persistence_json.cpp`, +92)

`\uXXXX` escapes are UTF-16 code units, so anything above the BMP arrives as a surrogate pair. These were encoded as two separate 3-byte sequences CESU-8, which is not valid UTF-8 - for keys and values alike. Now recombined into one 4-byte sequence, with unpaired surrogates rejected.


**Tests / perf / samples**

`test_layer_group_query_attention.cpp` (+93), `test_tokenizer.cpp` (+28), `test_tokenizer.py` (+42), `perf_layer.cpp` (+125: `Layer_GroupQueryAttention` MHA_ShortCache and Grouped_WithCache, plus `Layer_Sign`), and two samples (`granite_docling_inference.py`, `paddleocr_vl_inference.py`).


### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
2026-10-01 18:03:55 +03:00
Jaivardhan Bhola 84c2360c27 Merge pull request #29675 from jaivardhan-bhola:generalized-tokenizer
Generalize tokenizer loading to support method-based family dispatch in DNN - #29675

Companion PR: https://github.com/opencv/opencv_extra/pull/1402

### Changes
Added ALBERT and BERT support end-to-end, with `samples/dnn/albert_inference.py `and `samples/dnn/bert_inference.py `as validation samples, plus expanded coverage in modules/dnn/test/test_tokenizer.cpp. Required changes to `cv::dnn::dnn.hpp`, `graph_fusion_attention.cpp`, and `unicode.cpp/unicode.hpp` to support Unigram and WordPiece tokenizers.

To back ALBERT/BERT, generalized `cv::dnn::Tokenizer `from a single BPE implementation into a method-dispatched frontend, adding `core_wordpiece.cpp/hpp` (WordPiece) and `core_unigram.cpp/hpp` (Unigram) as new backends. `tokenizer.cpp` now routes by method across BPE, Gemma,, SentencePiece, Unigram, and WordPiece behind one shared interface.

Tested against the following samples and the output matches to old tokenizer:

```
gpt2_inference.py
qwen_inference.py
gemma3_inference.py
```
GPT2:

```
Preparing GPT-2 model...
Inferencing GPT-2 model...
Hello, I'm a language model, not a programming language. I'm a language model. I'm a language model. I'm a language model. I'm a language model. I'm a
```

Gemma3:
```
Preparing Gemma3 model...
Prompt:
<start_of_turn>user
What is OpenCV?<end_of_turn>
<start_of_turn>model

Inferencing Gemma3 model...
Response:
Okay, let's break down what OpenCV is.

**What is OpenCV?**

OpenCV (Open Source Computer Vision Library) is a powerful and
```

Qwen2.5:
```
Preparing Qwen2.5 model...
Prompt:
<|im_start|>user
What is OpenCV?<|im_end|>
<|im_start|>assistant

Inferencing Qwen2.5 model...
Response:
OpenCV is a set of computer vision libraries in C++ designed to be used for image and video processing. It provides a wide range of tools and functions for
```

### Tokenizer References 
Byte-level BPE: [tokenizers/src/pre_tokenizers/byte_level.rs](https://github.com/huggingface/tokenizers/blob/main/tokenizers/src/pre_tokenizers/byte_level.rs)
(This defines the byte-level mapping rules, which is used in conjunction with the [BPE model](https://www.google.com/search?q=https://github.com/huggingface/tokenizers/blob/main/tokenizers/src/models/bpe/mod.rs))

SentencePiece BPE (Metaspace): [tokenizers/src/pre_tokenizers/metaspace.rs](https://www.google.com/search?q=https://github.com/huggingface/tokenizers/blob/main/tokenizers/src/pre_tokenizers/metaspace.rs)
(This defines the rule for replacing whitespace with the U+2581 _ character and handling byte fallback)

Unigram: [tokenizers/src/models/unigram/mod.rs](https://www.google.com/search?q=https://github.com/huggingface/tokenizers/blob/main/tokenizers/src/models/unigram/mod.rs)
(This contains the core logic for the Unigram lattice scoring and probabilistic tokenization rules)

WordPiece: [tokenizers/src/models/wordpiece/mod.rs](https://www.google.com/search?q=https://github.com/huggingface/tokenizers/blob/main/tokenizers/src/models/wordpiece/mod.rs)
(This explicitly cites Schuster & Nakajima in the code comments and implements the greedy longest-match rule with the ## prefix)

### Pull Request Readiness Checklist

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on
      code under GPL or another license incompatible with OpenCV.
- [x] The PR is proposed to the proper branch (`5.x`).
- [x] There is a reference to the original bug report and related work.
- [x] There is accuracy test and test data in `opencv_extra`, same
      branch name (`generalized-tokenizer`) — `bert/`, `t5/` fixtures
      back the new C++ tests.
- [x] The feature is documented and sample code builds with project CMake.
2026-09-23 16:20:33 +03:00
Jorge Velez a49a293d3c Merge pull request #27534 from JorgeV92:gsoc2025-tokenizer
GSoC 2025: Add Tokenizer Support to DNN Module #27534

merge with https://github.com/opencv/opencv_extra/pull/1276

### Summary
This pull request introduces initial support for a tokenizer module under `modules/dnn/src/tokenizer` as part of Google Summer of Code 2025 (Project: Tokenization for OpenCV DNN).

### Status
- [x] Project structure in place
- [x] Initial BPE tokenizer loading
- [x] Regex splitting (in progress)
- [x] Encoding logic for GPT-2 tokenizer (in progress)
- [ ] Documentation (to be improved)

### Goals
The goal is to support Hugging Face-compatible tokenization (e.g., GPT-2) natively in C++ to be integrated with DNN inference pipelines. 

The core pipeline lives in `dnn/src/tokenizer/core_bpe.hpp` and `dnn/src/tokenizer/encoding.hpp`. For Unicode handling I’m using `dnn/src/tokenizer/unicode.hpp`, which is adapted from llama.cpp.


### Feedback
Please share early feedback on:
- General design structure
- Integration strategy with `dnn`
- Code organization or naming conventions

### Reference
Project: https://summerofcode.withgoogle.com/programs/2025/projects/79SW6eNK
2026-04-06 10:46:13 +03:00