Multi-Modal Multi-Granularity Tokenizer for Chu Bamboo Slip Scripts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yingfa, Hu, Chenlong, Feng, Cong, Song, Chenyang, Yu, Shi, Han, Xu, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
von: Song, Chenyang, et al.
Veröffentlicht: (2025)
von: Song, Chenyang, et al.
Veröffentlicht: (2025)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
von: Cao, Guiming, et al.
Veröffentlicht: (2024)
von: Cao, Guiming, et al.
Veröffentlicht: (2024)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding
von: Li, Zizhong, et al.
Veröffentlicht: (2026)
von: Li, Zizhong, et al.
Veröffentlicht: (2026)
$\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens
von: Zhang, Xinrong, et al.
Veröffentlicht: (2024)
von: Zhang, Xinrong, et al.
Veröffentlicht: (2024)
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
von: Zhang, Xinrong, et al.
Veröffentlicht: (2024)
von: Zhang, Xinrong, et al.
Veröffentlicht: (2024)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
von: Yu, Shi, et al.
Veröffentlicht: (2024)
von: Yu, Shi, et al.
Veröffentlicht: (2024)
StateX: Enhancing RNN Recall via Post-training State Expansion
von: Shen, Xingyu, et al.
Veröffentlicht: (2025)
von: Shen, Xingyu, et al.
Veröffentlicht: (2025)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
MMGR: Multi-Modal Generative Reasoning
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
von: Meituan LongCat Team, et al.
Veröffentlicht: (2026)
von: Meituan LongCat Team, et al.
Veröffentlicht: (2026)
Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion
von: Shi, Yi, et al.
Veröffentlicht: (2025)
von: Shi, Yi, et al.
Veröffentlicht: (2025)
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
von: Li, You, et al.
Veröffentlicht: (2025)
von: Li, You, et al.
Veröffentlicht: (2025)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
von: Xu, Runsen, et al.
Veröffentlicht: (2025)
von: Xu, Runsen, et al.
Veröffentlicht: (2025)
MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
Robust and Scalable Model Editing for Large Language Models
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
von: Yu, Tianyu, et al.
Veröffentlicht: (2023)
von: Yu, Tianyu, et al.
Veröffentlicht: (2023)
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
von: Yang, Honglong, et al.
Veröffentlicht: (2025)
von: Yang, Honglong, et al.
Veröffentlicht: (2025)
DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset
von: Lee, Young-Jun, et al.
Veröffentlicht: (2022)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2022)
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
Text-centric Alignment for Multi-Modality Learning
von: Tsai, Yun-Da, et al.
Veröffentlicht: (2024)
von: Tsai, Yun-Da, et al.
Veröffentlicht: (2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
Prompt Highlighter: Interactive Control for Multi-Modal LLMs
von: Zhang, Yuechen, et al.
Veröffentlicht: (2023)
von: Zhang, Yuechen, et al.
Veröffentlicht: (2023)
Shared and Private Information Learning in Multimodal Sentiment Analysis with Deep Modal Alignment and Self-supervised Multi-Task Learning
von: Lai, Songning, et al.
Veröffentlicht: (2023)
von: Lai, Songning, et al.
Veröffentlicht: (2023)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2024)
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2024)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
Core Knowledge Deficits in Multi-Modal Language Models
von: Li, Yijiang, et al.
Veröffentlicht: (2024)
von: Li, Yijiang, et al.
Veröffentlicht: (2024)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
Otter: A Multi-Modal Model with In-Context Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2023)
von: Li, Bo, et al.
Veröffentlicht: (2023)
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
von: Hu, Jinyi, et al.
Veröffentlicht: (2023)
von: Hu, Jinyi, et al.
Veröffentlicht: (2023)
Revisiting Multi-Modal LLM Evaluation
von: Lu, Jian, et al.
Veröffentlicht: (2024)
von: Lu, Jian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
von: Song, Chenyang, et al.
Veröffentlicht: (2025) -
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
von: Sun, Yubo, et al.
Veröffentlicht: (2025) -
Stuffed Mamba: Oversized States Lead to the Inability to Forget
von: Chen, Yingfa, et al.
Veröffentlicht: (2024) -
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
von: Cao, Guiming, et al.
Veröffentlicht: (2024) -
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)