SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Yiming, Wang, Junjie, Meng, Yuxin, Shi, Yihang, Lin, Zhiqiang, Chu, Ruihang, Xu, Yiran, Li, Ziming, Zhao, Yunfei, Wang, Zihan, Qiao, Yu, Tang, Ruiming, Liu, Minghao, Yang, Yujiu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
von: Wang, Junjie, et al.
Veröffentlicht: (2024)
von: Wang, Junjie, et al.
Veröffentlicht: (2024)
A 3D Framework for Improving Low-Latency Multi-Channel Live Streaming
von: Aiersilan, Aizierjiang, et al.
Veröffentlicht: (2024)
von: Aiersilan, Aizierjiang, et al.
Veröffentlicht: (2024)
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
von: Chen, Qi, et al.
Veröffentlicht: (2025)
von: Chen, Qi, et al.
Veröffentlicht: (2025)
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
von: Gao, Lancheng, et al.
Veröffentlicht: (2025)
von: Gao, Lancheng, et al.
Veröffentlicht: (2025)
SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in Conversation
von: Yi, Zijian, et al.
Veröffentlicht: (2024)
von: Yi, Zijian, et al.
Veröffentlicht: (2024)
SpikEmo: Enhancing Emotion Recognition With Spiking Temporal Dynamics in Conversations
von: Yu, Xiaomin, et al.
Veröffentlicht: (2024)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2024)
Muse: A Multimodal Conversational Recommendation Dataset with Scenario-Grounded User Profiles
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
Learning Video Context as Interleaved Multimodal Sequences
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
von: Zhao, Lingzhi, et al.
Veröffentlicht: (2025)
von: Zhao, Lingzhi, et al.
Veröffentlicht: (2025)
From Natural Alignment to Conditional Controllability in Multimodal Dialogue
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
Nagare Media Engine: A System for Cloud- and Edge-Native Network-based Multimedia Workflows
von: Neugebauer, Matthias
Veröffentlicht: (2025)
von: Neugebauer, Matthias
Veröffentlicht: (2025)
MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production
von: Hu, Huanran, et al.
Veröffentlicht: (2026)
von: Hu, Huanran, et al.
Veröffentlicht: (2026)
Apollo: Unified Multi-Task Audio-Video Joint Generation
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
Vidformer: Drop-in Declarative Optimization for Rendering Video-Native Query Results
von: Winecki, Dominik, et al.
Veröffentlicht: (2026)
von: Winecki, Dominik, et al.
Veröffentlicht: (2026)
Towards Open-Vocabulary Audio-Visual Event Localization
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
von: Lu, Jinghui, et al.
Veröffentlicht: (2024)
von: Lu, Jinghui, et al.
Veröffentlicht: (2024)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
von: Deng, Jingyuan, et al.
Veröffentlicht: (2025)
von: Deng, Jingyuan, et al.
Veröffentlicht: (2025)
Semantic-Guided Unsupervised Video Summarization
von: Liu, Haizhou, et al.
Veröffentlicht: (2026)
von: Liu, Haizhou, et al.
Veröffentlicht: (2026)
HCVR Scene Generation: High Compatibility Virtual Reality Environment Generation for Extended Redirected Walking
von: Zhang, Yiran, et al.
Veröffentlicht: (2026)
von: Zhang, Yiran, et al.
Veröffentlicht: (2026)
TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis
von: Quan, Weize, et al.
Veröffentlicht: (2024)
von: Quan, Weize, et al.
Veröffentlicht: (2024)
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
von: Wang, Xiaohan, et al.
Veröffentlicht: (2026)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2026)
Plasticity-Aware Mixture of Experts for Learning Under QoE Shifts in Adaptive Video Streaming
von: He, Zhiqiang, et al.
Veröffentlicht: (2025)
von: He, Zhiqiang, et al.
Veröffentlicht: (2025)
Accelerating Multi-Condition T2I Generation via Adaptive Condition Offloading and Pruning
von: Kong, Yuxin, et al.
Veröffentlicht: (2026)
von: Kong, Yuxin, et al.
Veröffentlicht: (2026)
Advancing Weakly-Supervised Audio-Visual Video Parsing via Segment-wise Pseudo Labeling
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
von: Pan, Zihan, et al.
Veröffentlicht: (2025)
von: Pan, Zihan, et al.
Veröffentlicht: (2025)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
von: Yang, Jialiang, et al.
Veröffentlicht: (2026)
von: Yang, Jialiang, et al.
Veröffentlicht: (2026)
SCI-Reason: A Dataset with Chain-of-Thought Rationales for Complex Multimodal Reasoning in Academic Areas
von: Ma, Chenghao, et al.
Veröffentlicht: (2025)
von: Ma, Chenghao, et al.
Veröffentlicht: (2025)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
Small Stickers, Big Meanings: A Multilingual Sticker Semantic Understanding Dataset with a Gamified Approach
von: Chee, Heng Er Metilda, et al.
Veröffentlicht: (2025)
von: Chee, Heng Er Metilda, et al.
Veröffentlicht: (2025)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
Personalized Playback Technology: How Short Video Services Create Excellent User Experience
von: Deng, Weihui, et al.
Veröffentlicht: (2024)
von: Deng, Weihui, et al.
Veröffentlicht: (2024)
TAVGBench: Benchmarking Text to Audible-Video Generation
von: Mao, Yuxin, et al.
Veröffentlicht: (2024)
von: Mao, Yuxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
von: Wang, Junjie, et al.
Veröffentlicht: (2024) -
A 3D Framework for Improving Low-Latency Multi-Channel Live Streaming
von: Aiersilan, Aizierjiang, et al.
Veröffentlicht: (2024) -
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
von: Chen, Qi, et al.
Veröffentlicht: (2025) -
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
von: Gao, Lancheng, et al.
Veröffentlicht: (2025) -
SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)