NVIDIA Developer
developer.nvidia.com › blog › robust-scene-text-detection-and-recognition-implementation
Robust Scene Text Detection and Recognition: Implementation | NVIDIA Technical Blog
November 14, 2024 - The scene text detection and recognition (STDR) pipeline is implemented using state-of-the-art deep learning models, including CRAFT for text detection and PARSeq for text recognition, and is optimized for accuracy and low latency.
Readthedocs
mmocr.readthedocs.io › en › latest › textrecog_models.html
Text Recognition Models — MMOCR 1.0.1 documentation
Correspondingly, we propose an autonomous, bidirectional and iterative ABINet for scene text recognition. Firstly, the autonomous suggests to block gradient flow between vision and language models to enforce explicitly language modeling. Secondly, a novel bidirectional cloze network (BCN) as the language model is proposed based on bidirectional feature representation.
[P]Modern open-source OCR capabilities and which model to choose
Look into open-mmlab's MMOCR, does both detection and recognition, with English and Chinese alphabet support. Absolutely wicked performance, it scrapes off text from logos, flyers, blurred text, etc. Not suitable for real-time performance. Until a few years ago, I was quite happy with Tesseract, but they've fallen behind since then. Still good for scanning printed text or similar. Also supports a lot of languages. More on reddit.com
Synology Photos Object Recognition is BACK!
Actually you don't need to manually install it. The Photos package will auto update in the usual phased roll-out for all packages. More on reddit.com
06:23
Python Machine Learning Project - Scene text recognition using ...
10:11
How to use OCR | Get Started with Optical Character Recognition ...
15:39
Text detection with Python and Opencv | OCR using EasyOCR | Computer ...
Accurate Text Detector Tutorial - Fast & Easy - Using EAST
05:29
Scene Text Image Super-Resolution Based on Text-Conditional Diffusion ...
52:42
[Open DMQA Seminar] Self- Semi-supervised Learning for Scene Text ...
arXiv
arxiv.org › abs › 2205.00159
[2205.00159] SVTR: Scene Text Recognition with a Single Visual Model
May 23, 2022 - Abstract:Dominant scene text recognition models commonly contain two building blocks, a visual model for feature extraction and a sequence model for text transcription. This hybrid architecture, although accurate, is complex and less efficient.
arXiv
arxiv.org › abs › 2306.16707
[2306.16707] DiffusionSTR: Diffusion Model for Scene Text Recognition
June 29, 2023 - Abstract:This paper presents Diffusion Model for Scene Text Recognition (DiffusionSTR), an end-to-end text recognition framework using diffusion models for recognizing text in the wild.
PubMed Central
pmc.ncbi.nlm.nih.gov › articles › PMC10181526
Lightweight Scene Text Recognition Based on Transformer - PMC
In this work, we balanced recognition accuracy, model speed, and cost. Based on Vit-tiny, we built a lightweight scene text recognition model (LSTR) and designed position-enhancement and visual-enhancement modules, allowing the model to maintain a fast inference speed while achieving higher accuracy at a lower cost.
GitHub
github.com › clovaai › deep-text-recognition-benchmark
GitHub - clovaai/deep-text-recognition-benchmark: Text recognition (optical character recognition) with deep learning methods, ICCV 2019 · GitHub
Official PyTorch implementation ... STR models fit into. Using this framework allows for the module-wise contributions to performance in terms of accuracy, speed, and memory demand, under one consistent set of training and evaluation datasets. Such analyses clean up the hindrance on the current comparisons to understand the performance gain of the existing modules. Based on this framework, we recorded the 1st place of ICDAR2013 focused scene text, ICDAR2019 ...
Starred by 3.9K users
Forked by 1.1K users
Languages Jupyter Notebook 90.4% | Python 9.6%
arXiv
arxiv.org › abs › 2302.14338
[2302.14338] Turning a CLIP Model into a Scene Text Detector
March 26, 2023 - The recent large-scale Contrastive Language-Image Pretraining (CLIP) model has shown great potential in various downstream tasks via leveraging the pretrained vision and language knowledge.
ArcGIS
doc.arcgis.com › en › pretrained-models › latest › imagery › introduction-to-scene-text-parsing.htm
Scene Text Parsing - Esri Documentation - ArcGIS Online
Extracting this text can provide additional context and details about the places the text describes and the information it conveys. This deep learning model is based on the PaddleOCR model and uses optical character recognition (OCR) technology to detect text in images.
Papers with Code
paperswithcode.com › task › scene-text-detection
Trending Papers - Hugging Face
EverMemOS presents a self-organizing memory system for large language models that processes dialogue streams into structured memory cells and scenes to enhance long-term interaction capabilities.
MDPI
mdpi.com › 2078-2489 › 14 › 7 › 369
Scene Text Recognition Based on Improved CRNN
June 28, 2023 - Section 2 summarizes the relevant research in this field, with a focus on text recognition in the field of deep learning. In Section 3, the CRNN recognition model, which combines label smoothing and the language model, is introduced in detail. Experimental results and corresponding discussions are provided in Section 4. The paper concludes with a summary in Section 5. As the field of deep learning continues to develop, the direction of scene text recognition has also been developed, and many researchers have proposed many new and relevant recognition algorithms.
GitHub
github.com › HCIILAB › Scene-Text-Recognition-Recommendations
GitHub - HCIILAB/Scene-Text-Recognition-Recommendations: Papers, Datasets, Algorithms, SOTA for STR. Long-time Maintaining · GitHub
arXiv-2023:Improving Scene Text Recognition for Character-Level Long-Tailed Distribution ... Neurocomputing-2023:DPF-S2S: A novel dual-pathway-fusion-based sequence-to-sequence text recognition model
Starred by 354 users
Forked by 39 users
Languages Python 96.7% | Shell 3.3%
TheCVF
openaccess.thecvf.com › content › CVPR2024 › papers › Xu_OTE_Exploring_Accurate_Scene_Text_Recognition_Using_One_Token_CVPR_2024_paper.pdf pdf
OTE: Exploring Accurate Scene Text Recognition Using One Token
MLT-5k [26] to train the model, which follows the setup of · [38]. The experiments are conducted on 4 NVIDIA 4090 · GPUs with batch size 384 per GPU for 20 epochs. Spe- cially, we directly use the scene text detector provided by · Wang et al. [38] without fine-tuning to crop the text patches. 4.3. Evaluation Metric · We set the size of the recognition character to 36, including ·