EUVD-2026-75629
Tesseract is an open source OCR engine. In version 5.5.3 and earlier, UNICHARSET::load_via_fgets in src/ccutil/unicharset.cpp trusts the declared unichar count as a loop bound and uses id as an unchecked index into the unichars vector. unichar_insert_backwards_compatible can leave the vector unchanged for an empty, duplicate, or already-encodable representation, causing id to become larger than unichars.size(). Subsequent set_* calls and the write to unichars[id].properties.enabled then write UNICHAR_PROPERTIES beyond the vector during initialization in both the default LSTM and legacy engines, causing heap corruption, a crash, or potentially controlled corruption. No fixed release is available as of this review.
- EUVD state
- Present in the current official mapping
- Known exploitation
- Not present in the current ENISA EUVD known-exploited dataset. This is not proof of no exploitation.
- ENISA score
- 7.8 · CVSS 3.1
- Advisory evidence
- No linked advisory details stored yet
