LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Danlu, Shi, Freda, Agarwal, Aditi, Myerston, Jacobo, Berg-Kirkpatrick, Taylor |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
por: Gao, Xin, et al.
Publicado: (2026)
por: Gao, Xin, et al.
Publicado: (2026)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
por: Shi, Chufan, et al.
Publicado: (2026)
por: Shi, Chufan, et al.
Publicado: (2026)
Learning Language Structures through Grounding
por: Shi, Freda
Publicado: (2024)
por: Shi, Freda
Publicado: (2024)
Optical Context Compression Is Just (Bad) Autoencoding
por: Lee, Ivan Yee, et al.
Publicado: (2025)
por: Lee, Ivan Yee, et al.
Publicado: (2025)
Alt-Text with Context: Improving Accessibility for Images on Twitter
por: Srivatsan, Nikita, et al.
Publicado: (2023)
por: Srivatsan, Nikita, et al.
Publicado: (2023)
Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages
por: Chen, Danlu, et al.
Publicado: (2026)
por: Chen, Danlu, et al.
Publicado: (2026)
CLEAR: Character Unlearning in Textual and Visual Modalities
por: Dontsov, Alexey, et al.
Publicado: (2024)
por: Dontsov, Alexey, et al.
Publicado: (2024)
Enhancing Steganographic Text Extraction: Evaluating the Impact of NLP Models on Accuracy and Semantic Coherence
por: Li, Mingyang, et al.
Publicado: (2024)
por: Li, Mingyang, et al.
Publicado: (2024)
Pixel Sentence Representation Learning
por: Xiao, Chenghao, et al.
Publicado: (2024)
por: Xiao, Chenghao, et al.
Publicado: (2024)
What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric
por: Kerkouri, Mohamed Amine, et al.
Publicado: (2026)
por: Kerkouri, Mohamed Amine, et al.
Publicado: (2026)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
por: Jiang, Yifan, et al.
Publicado: (2026)
por: Jiang, Yifan, et al.
Publicado: (2026)
Beyond the Textual: Generating Coherent Visual Options for MCQs
por: Wang, Wanqiang, et al.
Publicado: (2025)
por: Wang, Wanqiang, et al.
Publicado: (2025)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
por: Gordon, Brian, et al.
Publicado: (2023)
por: Gordon, Brian, et al.
Publicado: (2023)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
por: Yang, Cheng, et al.
Publicado: (2026)
por: Yang, Cheng, et al.
Publicado: (2026)
Tell Me What's Next: Textual Foresight for Generic UI Representations
por: Burns, Andrea, et al.
Publicado: (2024)
por: Burns, Andrea, et al.
Publicado: (2024)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
por: Hashemi, Mohammad Abuzar, et al.
Publicado: (2021)
por: Hashemi, Mohammad Abuzar, et al.
Publicado: (2021)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
por: Maini, Pratyush, et al.
Publicado: (2023)
por: Maini, Pratyush, et al.
Publicado: (2023)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
por: Sun, Hao, et al.
Publicado: (2026)
por: Sun, Hao, et al.
Publicado: (2026)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
por: Wang, Yifan, et al.
Publicado: (2026)
por: Wang, Yifan, et al.
Publicado: (2026)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
por: Dhawan, Aashish, et al.
Publicado: (2026)
por: Dhawan, Aashish, et al.
Publicado: (2026)
torchdistill Meets Hugging Face Libraries for Reproducible, Coding-Free Deep Learning Studies: A Case Study on NLP
por: Matsubara, Yoshitomo
Publicado: (2023)
por: Matsubara, Yoshitomo
Publicado: (2023)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
por: Hua, Jiacheng, et al.
Publicado: (2026)
por: Hua, Jiacheng, et al.
Publicado: (2026)
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
por: Zhang, Zheyuan, et al.
Publicado: (2024)
por: Zhang, Zheyuan, et al.
Publicado: (2024)
FSMR: A Feature Swapping Multi-modal Reasoning Approach with Joint Textual and Visual Clues
por: Li, Shuang, et al.
Publicado: (2024)
por: Li, Shuang, et al.
Publicado: (2024)
The Mechanistic Emergence of Symbol Grounding in Language Models
por: Wu, Shuyu, et al.
Publicado: (2025)
por: Wu, Shuyu, et al.
Publicado: (2025)
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
por: Wang, Chenglong, et al.
Publicado: (2024)
por: Wang, Chenglong, et al.
Publicado: (2024)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
por: Zhang, Beichen, et al.
Publicado: (2025)
por: Zhang, Beichen, et al.
Publicado: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
por: Guo, Ziyu, et al.
Publicado: (2025)
por: Guo, Ziyu, et al.
Publicado: (2025)
VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites
por: Islam, Md. Adnanul, et al.
Publicado: (2025)
por: Islam, Md. Adnanul, et al.
Publicado: (2025)
Visual Representations inside the Language Model
por: Liu, Benlin, et al.
Publicado: (2025)
por: Liu, Benlin, et al.
Publicado: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
por: Wu, Juncheng, et al.
Publicado: (2026)
por: Wu, Juncheng, et al.
Publicado: (2026)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
por: Ge, Jinchao, et al.
Publicado: (2025)
por: Ge, Jinchao, et al.
Publicado: (2025)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
por: Gan, Woody Haosheng, et al.
Publicado: (2025)
por: Gan, Woody Haosheng, et al.
Publicado: (2025)
Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting
por: Kayser, Maxime, et al.
Publicado: (2024)
por: Kayser, Maxime, et al.
Publicado: (2024)
Emergent Visual-Semantic Hierarchies in Image-Text Representations
por: Alper, Morris, et al.
Publicado: (2024)
por: Alper, Morris, et al.
Publicado: (2024)
EVA-02: A Visual Representation for Neon Genesis
por: Fang, Yuxin, et al.
Publicado: (2023)
por: Fang, Yuxin, et al.
Publicado: (2023)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
por: Zhu, Wang, et al.
Publicado: (2023)
por: Zhu, Wang, et al.
Publicado: (2023)
A Review on Large Language Models for Visual Analytics
por: Agarwal, Navya Sonal, et al.
Publicado: (2025)
por: Agarwal, Navya Sonal, et al.
Publicado: (2025)
PLVS: A SLAM System with Points, Lines, Volumetric Mapping, and 3D Incremental Segmentation
por: Freda, Luigi
Publicado: (2023)
por: Freda, Luigi
Publicado: (2023)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
por: Verma, Gaurav, et al.
Publicado: (2024)
por: Verma, Gaurav, et al.
Publicado: (2024)
Ejemplares similares
-
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
por: Gao, Xin, et al.
Publicado: (2026) -
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
por: Shi, Chufan, et al.
Publicado: (2026) -
Learning Language Structures through Grounding
por: Shi, Freda
Publicado: (2024) -
Optical Context Compression Is Just (Bad) Autoencoding
por: Lee, Ivan Yee, et al.
Publicado: (2025) -
Alt-Text with Context: Improving Accessibility for Images on Twitter
por: Srivatsan, Nikita, et al.
Publicado: (2023)