ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Feng, Shi, Zesheng, Wang, Bo, Wang, Nan, Xiao, Han |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024)
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024)
TriAlignGR: Triangular Multitask Alignment with Multimodal Deep Interest Mining for Generative Recommendation
von: Zeng, Yangchen, et al.
Veröffentlicht: (2026)
von: Zeng, Yangchen, et al.
Veröffentlicht: (2026)
jina-reranker-v3: Last but Not Late Interaction for Listwise Document Reranking
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
von: Sturua, Saba, et al.
Veröffentlicht: (2024)
von: Sturua, Saba, et al.
Veröffentlicht: (2024)
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
von: Günther, Michael, et al.
Veröffentlicht: (2024)
von: Günther, Michael, et al.
Veröffentlicht: (2024)
Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
von: Jha, Rohan, et al.
Veröffentlicht: (2024)
von: Jha, Rohan, et al.
Veröffentlicht: (2024)
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
von: Günther, Michael, et al.
Veröffentlicht: (2025)
von: Günther, Michael, et al.
Veröffentlicht: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
von: Farzulla, Murad
Veröffentlicht: (2026)
von: Farzulla, Murad
Veröffentlicht: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
The MSR-Video to Text Dataset with Clean Annotations
von: Chen, Haoran, et al.
Veröffentlicht: (2021)
von: Chen, Haoran, et al.
Veröffentlicht: (2021)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
von: Skripkin, Matvey, et al.
Veröffentlicht: (2025)
von: Skripkin, Matvey, et al.
Veröffentlicht: (2025)
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
von: Mohr, Isabelle, et al.
Veröffentlicht: (2024)
von: Mohr, Isabelle, et al.
Veröffentlicht: (2024)
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
A Study into Investigating Temporal Robustness of LLMs
von: Wallat, Jonas, et al.
Veröffentlicht: (2025)
von: Wallat, Jonas, et al.
Veröffentlicht: (2025)
Evaluation of Table Representations to Answer Questions from Tables in Documents : A Case Study using 3GPP Specifications
von: Roychowdhury, Sujoy, et al.
Veröffentlicht: (2024)
von: Roychowdhury, Sujoy, et al.
Veröffentlicht: (2024)
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
von: Bian, Zhipeng, et al.
Veröffentlicht: (2025)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2025)
PairCFR: Enhancing Model Training on Paired Counterfactually Augmented Data through Contrastive Learning
von: Qiu, Xiaoqi, et al.
Veröffentlicht: (2024)
von: Qiu, Xiaoqi, et al.
Veröffentlicht: (2024)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025)
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
PDFMathTranslate: Scientific Document Translation Preserving Layouts
von: Ouyang, Rongxin, et al.
Veröffentlicht: (2025)
von: Ouyang, Rongxin, et al.
Veröffentlicht: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
von: Sáez, Arnau Igualde, et al.
Veröffentlicht: (2025)
von: Sáez, Arnau Igualde, et al.
Veröffentlicht: (2025)
A Survey on Natural Language Counterfactual Generation
von: Wang, Yongjie, et al.
Veröffentlicht: (2024)
von: Wang, Yongjie, et al.
Veröffentlicht: (2024)
Efficient Code Embeddings from Code Generation Models
von: Kryvosheieva, Daria, et al.
Veröffentlicht: (2025)
von: Kryvosheieva, Daria, et al.
Veröffentlicht: (2025)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
von: Qi, Dekang, et al.
Veröffentlicht: (2026)
von: Qi, Dekang, et al.
Veröffentlicht: (2026)
IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
RONA: Pragmatically Diverse Image Captioning with Coherence Relations
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
XPath Agent: An Efficient XPath Programming Agent Based on LLM for Web Crawler
von: Li, Yu, et al.
Veröffentlicht: (2024)
von: Li, Yu, et al.
Veröffentlicht: (2024)
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
Empowering Tabular Data Preparation with Language Models: Why and How?
von: Chen, Mengshi, et al.
Veröffentlicht: (2025)
von: Chen, Mengshi, et al.
Veröffentlicht: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
von: Fan, Jingxing, et al.
Veröffentlicht: (2025)
von: Fan, Jingxing, et al.
Veröffentlicht: (2025)
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
von: Vileikytė, Brigita, et al.
Veröffentlicht: (2024)
von: Vileikytė, Brigita, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024) -
TriAlignGR: Triangular Multitask Alignment with Multimodal Deep Interest Mining for Generative Recommendation
von: Zeng, Yangchen, et al.
Veröffentlicht: (2026) -
jina-reranker-v3: Last but Not Late Interaction for Listwise Document Reranking
von: Wang, Feng, et al.
Veröffentlicht: (2025) -
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
von: Sturua, Saba, et al.
Veröffentlicht: (2024) -
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
von: Günther, Michael, et al.
Veröffentlicht: (2024)