Lumos : Empowering Multimodal LLMs with Scene Text Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Shenoy, Ashish, Lu, Yichao, Jayakumar, Srihari, Chatterjee, Debojeet, Moslehpour, Mohsen, Chuang, Pierce, Harpale, Abhay, Bhardwaj, Vikas, Xu, Di, Zhao, Shicong, Zhao, Longfang, Ramchandani, Ankit, Dong, Xin Luna, Kumar, Anuj |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EgoQR: Efficient QR Code Reading in Egocentric Settings
por: Moslehpour, Mohsen, et al.
Publicado: (2024)
por: Moslehpour, Mohsen, et al.
Publicado: (2024)
GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables
por: Ramachandran, Akhil, et al.
Publicado: (2026)
por: Ramachandran, Akhil, et al.
Publicado: (2026)
Now It Sounds Like You: Learning Personalized Vocabulary On Device
por: Wang, Sid, et al.
Publicado: (2023)
por: Wang, Sid, et al.
Publicado: (2023)
Lumos Extrema
por: Moitra, Upamanyu
Publicado: (2024)
por: Moitra, Upamanyu
Publicado: (2024)
Historical and contemporary trends in competitive balance in the Commonwealth Games
por: Girish Ramchandani
Publicado: (2014)
por: Girish Ramchandani
Publicado: (2014)
Lumos3D: A Single-Forward Framework for Low-Light 3D Scene Restoration
por: Liu, Hanzhou, et al.
Publicado: (2025)
por: Liu, Hanzhou, et al.
Publicado: (2025)
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
por: Zhangli, Qilong, et al.
Publicado: (2024)
por: Zhangli, Qilong, et al.
Publicado: (2024)
Lumos: Let there be Language Model System Certification
por: Chaudhary, Isha, et al.
Publicado: (2025)
por: Chaudhary, Isha, et al.
Publicado: (2025)
UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
por: Liu, Ropeway, et al.
Publicado: (2025)
por: Liu, Ropeway, et al.
Publicado: (2025)
Shot-frugal and Robust quantum kernel classifiers
por: Shastry, Abhay, et al.
Publicado: (2022)
por: Shastry, Abhay, et al.
Publicado: (2022)
skotheimlab/xie_etal_2025_autonomous_cell_size_control: Pre-publication release
por: Shicong Xie
Publicado: (2025)
por: Shicong Xie
Publicado: (2025)
heyLeoLin/Adjacent-CRGs-mulstistep-deblending
por: Lin Shicong
Publicado: (2026)
por: Lin Shicong
Publicado: (2026)
Seismic Deblending via an Adjacent-CRG-Assisted Multistep Deep Learning Framework
por: Lin Shicong
Publicado: (2026)
por: Lin Shicong
Publicado: (2026)
A Miniature Vision-Based Localization System for Indoor Blimps
por: Ma, Shicong
Publicado: (2024)
por: Ma, Shicong
Publicado: (2024)
Ultrasound‐based radiomics for the differential diagnosis of breast masses: A systematic review and meta‐analysis
por: Xuerong Li, et al.
Publicado: (2024)
por: Xuerong Li, et al.
Publicado: (2024)
Driving Circular Economy Practices in Electronics and Electrical Industry in China: The Roles of Green Knowledge Sharing and Collaboration, and Digitalization
por: Wei Sun, et al.
Publicado: (2026)
por: Wei Sun, et al.
Publicado: (2026)
LumosFlow: Motion-Guided Long Video Generation
por: Chen, Jiahao, et al.
Publicado: (2025)
por: Chen, Jiahao, et al.
Publicado: (2025)
Root Causing Prediction Anomalies Using Explainable AI
por: Vishnampet, Ramanathan, et al.
Publicado: (2024)
por: Vishnampet, Ramanathan, et al.
Publicado: (2024)
Enhancing Li‐Ion Electric Vehicle Battery Performance: Analysis of Advanced Thermal Management Cooling Strategies with Phase Change Material
por: Ashish Dewangan, et al.
Publicado: (2024)
por: Ashish Dewangan, et al.
Publicado: (2024)
Exploring factors influencing the entrepreneurial intentions of the youth community towards green ICT to encourage environmental sustainability: Evidence from an emerging economy
por: Shivam Bhardwaj, et al.
Publicado: (2024)
por: Shivam Bhardwaj, et al.
Publicado: (2024)
Machine learning driven high-resolution Raman spectral generation for accurate molecular feature recognition
por: Yadav, Vikas, et al.
Publicado: (2024)
por: Yadav, Vikas, et al.
Publicado: (2024)
LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
por: Xing, Jiazheng, et al.
Publicado: (2026)
por: Xing, Jiazheng, et al.
Publicado: (2026)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
por: Yin, Da, et al.
Publicado: (2023)
por: Yin, Da, et al.
Publicado: (2023)
Lumos: Heterogeneity-aware Federated Graph Learning over Decentralized Devices
por: Pan, Qiying, et al.
Publicado: (2023)
por: Pan, Qiying, et al.
Publicado: (2023)
Wherefore Art Thou? Provenance-Guided Automatic Online Debugging with Lumos
por: Chen, Jingyuan, et al.
Publicado: (2026)
por: Chen, Jingyuan, et al.
Publicado: (2026)
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
por: Liang, Mingyu, et al.
Publicado: (2025)
por: Liang, Mingyu, et al.
Publicado: (2025)
Deep Learning Meets Mechanism Design: Key Results and Some Novel Applications
por: Sankar, V. Udaya, et al.
Publicado: (2024)
por: Sankar, V. Udaya, et al.
Publicado: (2024)
Hydrophobically modified nano‐SiO2 in situ loaded on polyurethane foam as a potential carrier for marine oil spill
por: Yan Hu, et al.
Publicado: (2024)
por: Yan Hu, et al.
Publicado: (2024)
SWAN: Self-supervised Wavelet Neural Network for Hyperspectral Image Unmixing
por: Ramchandani, Yassh, et al.
Publicado: (2025)
por: Ramchandani, Yassh, et al.
Publicado: (2025)
DepthPark : Smart, Cost‐Effective Vision‐Based Indoor Parking Management System Using Single Monocular Depth Estimation
por: Lakshay Naresh Ramchandani, et al.
Publicado: (2026)
por: Lakshay Naresh Ramchandani, et al.
Publicado: (2026)
In Silico Identification of TYK2 Pseudokinase Inhibitors Using Machine Learning and MD Simulations
por: Manish Ramchandani, et al.
Publicado: (2025)
por: Manish Ramchandani, et al.
Publicado: (2025)
FieldFormer: Locality-Aware Transformers for Spatio-Temporal Modeling on Sparse Sensor Networks
por: Bhardwaj, Ankit, et al.
Publicado: (2025)
por: Bhardwaj, Ankit, et al.
Publicado: (2025)
The social life of climate projects
por: Malcolm Araos, et al.
Publicado: (2024)
por: Malcolm Araos, et al.
Publicado: (2024)
Low-Complexity Near-Field Localization with XL-MIMO Sectored Uniform Circular Arrays
por: Liu, Shicong, et al.
Publicado: (2024)
por: Liu, Shicong, et al.
Publicado: (2024)
PhishLumos: An Adaptive Multi-Agent System for Proactive Phishing Campaign Mitigation
por: Chiba, Daiki, et al.
Publicado: (2025)
por: Chiba, Daiki, et al.
Publicado: (2025)
Deep Learning Based Auction Design for Selling Agricultural Produce through Farmer Collectives to Maximize Nash Social Welfare
por: Bhardwaj, Mayank Ratan, et al.
Publicado: (2025)
por: Bhardwaj, Mayank Ratan, et al.
Publicado: (2025)
Optimal binary codes from $\mathcal{C}_{D}$-codes over a non-chain ring
por: Yadav, Ankit, et al.
Publicado: (2025)
por: Yadav, Ankit, et al.
Publicado: (2025)
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
por: Chen, Zeyu, et al.
Publicado: (2026)
por: Chen, Zeyu, et al.
Publicado: (2026)
Photochemical Direct Mono/Di/Trifluoro‐Functionalization of Quinoxalin‐2(1 H )‐ones
por: Debabrat Goswami, et al.
Publicado: (2025)
por: Debabrat Goswami, et al.
Publicado: (2025)
Lumos: Performance Characterization of WebAssembly as a Serverless Runtime in the Edge-Cloud Continuum
por: Marcelino, Cynthia, et al.
Publicado: (2025)
por: Marcelino, Cynthia, et al.
Publicado: (2025)
Ejemplares similares
-
EgoQR: Efficient QR Code Reading in Egocentric Settings
por: Moslehpour, Mohsen, et al.
Publicado: (2024) -
GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables
por: Ramachandran, Akhil, et al.
Publicado: (2026) -
Now It Sounds Like You: Learning Personalized Vocabulary On Device
por: Wang, Sid, et al.
Publicado: (2023) -
Lumos Extrema
por: Moitra, Upamanyu
Publicado: (2024) -
Historical and contemporary trends in competitive balance in the Commonwealth Games
por: Girish Ramchandani
Publicado: (2014)