Token Entropy Regularization for Multi-modal Antenna Affiliation Identification
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Dong, Li, Ruoyu, Zhang, Xinyan, Xu, Jialei, Zhao, Ruosen, Zhang, Zhikang, Li, Lingyun, Wei, Zizhuang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
by: Zhao, Ruosen, et al.
Published: (2025)
by: Zhao, Ruosen, et al.
Published: (2025)
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
by: Gu, Bo, et al.
Published: (2026)
by: Gu, Bo, et al.
Published: (2026)
CitySeg: A 3D Open Vocabulary Semantic Segmentation Foundation Model in City-scale Scenarios
by: Xu, Jialei, et al.
Published: (2025)
by: Xu, Jialei, et al.
Published: (2025)
Clinical-Prior Guided Multi-Modal Learning with Latent Attention Pooling for Gait-Based Scoliosis Screening
by: Chen, Dong, et al.
Published: (2026)
by: Chen, Dong, et al.
Published: (2026)
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
by: Qiu, Xuerui, et al.
Published: (2026)
by: Qiu, Xuerui, et al.
Published: (2026)
Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval
by: Li, Wenjun, et al.
Published: (2024)
by: Li, Wenjun, et al.
Published: (2024)
Video Token Merging for Long-form Video Understanding
by: Lee, Seon-Ho, et al.
Published: (2024)
by: Lee, Seon-Ho, et al.
Published: (2024)
FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
by: Jing, Liqiang, et al.
Published: (2023)
by: Jing, Liqiang, et al.
Published: (2023)
UniVG: Towards UNIfied-modal Video Generation
by: Ruan, Ludan, et al.
Published: (2024)
by: Ruan, Ludan, et al.
Published: (2024)
Multi-modal Multi-platform Person Re-Identification: Benchmark and Method
by: Ha, Ruiyang, et al.
Published: (2025)
by: Ha, Ruiyang, et al.
Published: (2025)
Multi-modality Affinity Inference for Weakly Supervised 3D Semantic Segmentation
by: Li, Xiawei, et al.
Published: (2023)
by: Li, Xiawei, et al.
Published: (2023)
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy
by: Shen, Yunhang, et al.
Published: (2025)
by: Shen, Yunhang, et al.
Published: (2025)
Rényi Entropy: A New Token Pruning Metric for Vision Transformers
by: Su, Wei-Yuan, et al.
Published: (2026)
by: Su, Wei-Yuan, et al.
Published: (2026)
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Efficient Multi-modal Large Language Models via Visual Token Grouping
by: Huang, Minbin, et al.
Published: (2024)
by: Huang, Minbin, et al.
Published: (2024)
Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks
by: Yu, Kairong, et al.
Published: (2025)
by: Yu, Kairong, et al.
Published: (2025)
HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection
by: Cui, Jialei, et al.
Published: (2025)
by: Cui, Jialei, et al.
Published: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
by: Chen, Xinyan, et al.
Published: (2025)
by: Chen, Xinyan, et al.
Published: (2025)
OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
by: Yang, Morunliu, et al.
Published: (2026)
by: Yang, Morunliu, et al.
Published: (2026)
OmniColor: A Unified Framework for Multi-modal Lineart Colorization
by: Zhang, Xulu, et al.
Published: (2026)
by: Zhang, Xulu, et al.
Published: (2026)
MoST: Multi-modality Scene Tokenization for Motion Prediction
by: Mu, Norman, et al.
Published: (2024)
by: Mu, Norman, et al.
Published: (2024)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
Multi-modality action recognition based on dual feature shift in vehicle cabin monitoring
by: Lin, Dan, et al.
Published: (2024)
by: Lin, Dan, et al.
Published: (2024)
WhisperNet: A Scalable Solution for Bandwidth-Efficient Collaboration
by: Chen, Gong, et al.
Published: (2026)
by: Chen, Gong, et al.
Published: (2026)
TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
by: Hu, Junshan, et al.
Published: (2025)
by: Hu, Junshan, et al.
Published: (2025)
Empowering Segmentation Ability to Multi-modal Large Language Models
by: Yang, Yuqi, et al.
Published: (2024)
by: Yang, Yuqi, et al.
Published: (2024)
C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection
by: Wang, Siheng, et al.
Published: (2025)
by: Wang, Siheng, et al.
Published: (2025)
Phantom-Insight: Adaptive Multi-cue Fusion for Video Camouflaged Object Detection with Multimodal LLM
by: Zhang, Hua, et al.
Published: (2025)
by: Zhang, Hua, et al.
Published: (2025)
Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
by: Liu, Zhihang, et al.
Published: (2025)
by: Liu, Zhihang, et al.
Published: (2025)
Unified Multi-modal Diagnostic Framework with Reconstruction Pre-training and Heterogeneity-combat Tuning
by: Zhang, Yupei, et al.
Published: (2024)
by: Zhang, Yupei, et al.
Published: (2024)
SimDistill: Simulated Multi-modal Distillation for BEV 3D Object Detection
by: Zhao, Haimei, et al.
Published: (2023)
by: Zhao, Haimei, et al.
Published: (2023)
SDGE: Stereo Guided Depth Estimation for 360$^\circ$ Camera Sets
by: Xu, Jialei, et al.
Published: (2024)
by: Xu, Jialei, et al.
Published: (2024)
TokenTrace: Multi-Concept Attribution through Watermarked Token Recovery
by: Zhang, Li, et al.
Published: (2026)
by: Zhang, Li, et al.
Published: (2026)
EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization
by: Ji, Yingrui, et al.
Published: (2025)
by: Ji, Yingrui, et al.
Published: (2025)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
by: Zhang, Kaichen, et al.
Published: (2024)
by: Zhang, Kaichen, et al.
Published: (2024)
Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity Optimization
by: Chen, Jiayu, et al.
Published: (2026)
by: Chen, Jiayu, et al.
Published: (2026)
Similar Items
-
SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
by: Zhao, Ruosen, et al.
Published: (2025) -
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
by: Gu, Bo, et al.
Published: (2026) -
CitySeg: A 3D Open Vocabulary Semantic Segmentation Foundation Model in City-scale Scenarios
by: Xu, Jialei, et al.
Published: (2025) -
Clinical-Prior Guided Multi-Modal Learning with Latent Attention Pooling for Gait-Based Scoliosis Screening
by: Chen, Dong, et al.
Published: (2026) -
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
by: Zhang, Pingping, et al.
Published: (2024)