A Framework For Refining Text Classification and Object Recognition from Academic Articles
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Jinghong, Ota, Koichi, Gu, Wen, Hasegawa, Shinobu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Object Recognition from Scientific Document based on Compartment Refinement Framework
por: Li, Jinghong, et al.
Publicado: (2023)
por: Li, Jinghong, et al.
Publicado: (2023)
Hierarchical Tree-structured Knowledge Graph For Academic Insight Survey
por: Li, Jinghong, et al.
Publicado: (2024)
por: Li, Jinghong, et al.
Publicado: (2024)
A Survey Forest Diagram : Gain a Divergent Insight View on a Specific Research Topic
por: Li, Jinghong, et al.
Publicado: (2024)
por: Li, Jinghong, et al.
Publicado: (2024)
HATFormer: Historic Handwritten Arabic Text Recognition with Transformers
por: Chan, Adrian, et al.
Publicado: (2024)
por: Chan, Adrian, et al.
Publicado: (2024)
Lumos : Empowering Multimodal LLMs with Scene Text Recognition
por: Shenoy, Ashish, et al.
Publicado: (2024)
por: Shenoy, Ashish, et al.
Publicado: (2024)
Text Role Classification in Scientific Charts Using Multimodal Transformers
por: Kim, Hye Jin, et al.
Publicado: (2024)
por: Kim, Hye Jin, et al.
Publicado: (2024)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
por: Motamed, Saman, et al.
Publicado: (2023)
por: Motamed, Saman, et al.
Publicado: (2023)
JSTR: Judgment Improves Scene Text Recognition
por: Fujitake, Masato
Publicado: (2024)
por: Fujitake, Masato
Publicado: (2024)
Mostly Text, Smart Visuals: Asymmetric Text-Visual Pruning for Large Vision-Language Models
por: Li, Sijie, et al.
Publicado: (2026)
por: Li, Sijie, et al.
Publicado: (2026)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
por: Cheng, Jiale, et al.
Publicado: (2025)
por: Cheng, Jiale, et al.
Publicado: (2025)
Glyph: Scaling Context Windows via Visual-Text Compression
por: Cheng, Jiale, et al.
Publicado: (2025)
por: Cheng, Jiale, et al.
Publicado: (2025)
C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion
por: Yoon, Hee Suk, et al.
Publicado: (2024)
por: Yoon, Hee Suk, et al.
Publicado: (2024)
Text-centric Alignment for Multi-Modality Learning
por: Tsai, Yun-Da, et al.
Publicado: (2024)
por: Tsai, Yun-Da, et al.
Publicado: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
por: Salamatian, Ali, et al.
Publicado: (2025)
por: Salamatian, Ali, et al.
Publicado: (2025)
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
por: Mei, Jingbiao, et al.
Publicado: (2025)
por: Mei, Jingbiao, et al.
Publicado: (2025)
Text-guided Controllable Mesh Refinement for Interactive 3D Modeling
por: Chen, Yun-Chun, et al.
Publicado: (2024)
por: Chen, Yun-Chun, et al.
Publicado: (2024)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
por: Lei, Jiayi, et al.
Publicado: (2025)
por: Lei, Jiayi, et al.
Publicado: (2025)
Analytical Softmax Temperature Setting from Feature Dimensions for Model- and Domain-Robust Classification
por: Hasegawa, Tatsuhito, et al.
Publicado: (2025)
por: Hasegawa, Tatsuhito, et al.
Publicado: (2025)
DreamReward: Text-to-3D Generation with Human Preference
por: Ye, Junliang, et al.
Publicado: (2024)
por: Ye, Junliang, et al.
Publicado: (2024)
Indian Sign Language Recognition Using Mediapipe Holistic
por: G, Velmathi, et al.
Publicado: (2023)
por: G, Velmathi, et al.
Publicado: (2023)
Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
por: Park, Kwanyong, et al.
Publicado: (2024)
por: Park, Kwanyong, et al.
Publicado: (2024)
Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation
por: Chang, Wei-Chia, et al.
Publicado: (2025)
por: Chang, Wei-Chia, et al.
Publicado: (2025)
T$^3$Bench: Benchmarking Current Progress in Text-to-3D Generation
por: He, Yuze, et al.
Publicado: (2023)
por: He, Yuze, et al.
Publicado: (2023)
Calibrating Multimodal Consensus for Emotion Recognition
por: Zhong, Guowei, et al.
Publicado: (2025)
por: Zhong, Guowei, et al.
Publicado: (2025)
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models
por: Yi, Hao, et al.
Publicado: (2024)
por: Yi, Hao, et al.
Publicado: (2024)
SciGA: A Comprehensive Dataset for Designing Graphical Abstracts in Academic Papers
por: Kawada, Takuro, et al.
Publicado: (2025)
por: Kawada, Takuro, et al.
Publicado: (2025)
A Comparative Study of Machine Unlearning Techniques for Image and Text Classification Models
por: Safa, Omar M., et al.
Publicado: (2024)
por: Safa, Omar M., et al.
Publicado: (2024)
Generative Technology for Human Emotion Recognition: A Scope Review
por: Ma, Fei, et al.
Publicado: (2024)
por: Ma, Fei, et al.
Publicado: (2024)
Low-Resource Heuristics for Bahnaric Optical Character Recognition Improvement
por: Tran, Phat, et al.
Publicado: (2026)
por: Tran, Phat, et al.
Publicado: (2026)
What Shape Is Optimal for Masks in Text Removal?
por: Nakada, Hyakka, et al.
Publicado: (2025)
por: Nakada, Hyakka, et al.
Publicado: (2025)
Evaluating Numerical Reasoning in Text-to-Image Models
por: Kajić, Ivana, et al.
Publicado: (2024)
por: Kajić, Ivana, et al.
Publicado: (2024)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
por: Chen, Zhaorun, et al.
Publicado: (2024)
por: Chen, Zhaorun, et al.
Publicado: (2024)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
por: Zhou, Yiyang, et al.
Publicado: (2023)
por: Zhou, Yiyang, et al.
Publicado: (2023)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
por: Yuan, Huizhuo, et al.
Publicado: (2024)
por: Yuan, Huizhuo, et al.
Publicado: (2024)
Alt-Text with Context: Improving Accessibility for Images on Twitter
por: Srivatsan, Nikita, et al.
Publicado: (2023)
por: Srivatsan, Nikita, et al.
Publicado: (2023)
Generating Fine Details of Entity Interactions
por: Gu, Xinyi, et al.
Publicado: (2025)
por: Gu, Xinyi, et al.
Publicado: (2025)
Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance
por: Zhao, Linxi, et al.
Publicado: (2024)
por: Zhao, Linxi, et al.
Publicado: (2024)
SemIRNet: A Semantic Irony Recognition Network for Multimodal Sarcasm Detection
por: Zhou, Jingxuan, et al.
Publicado: (2025)
por: Zhou, Jingxuan, et al.
Publicado: (2025)
MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition
por: Fritsch, Stefan Gerd, et al.
Publicado: (2024)
por: Fritsch, Stefan Gerd, et al.
Publicado: (2024)
Ejemplares similares
-
Object Recognition from Scientific Document based on Compartment Refinement Framework
por: Li, Jinghong, et al.
Publicado: (2023) -
Hierarchical Tree-structured Knowledge Graph For Academic Insight Survey
por: Li, Jinghong, et al.
Publicado: (2024) -
A Survey Forest Diagram : Gain a Divergent Insight View on a Specific Research Topic
por: Li, Jinghong, et al.
Publicado: (2024) -
HATFormer: Historic Handwritten Arabic Text Recognition with Transformers
por: Chan, Adrian, et al.
Publicado: (2024) -
Lumos : Empowering Multimodal LLMs with Scene Text Recognition
por: Shenoy, Ashish, et al.
Publicado: (2024)