Foundations of Multisensory Artificial Intelligence
Fuente:
arXiv
Saved in:
| Main Author: | Liang, Paul Pu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and Reasoning
by: Zhou, Chuhao, et al.
Published: (2025)
by: Zhou, Chuhao, et al.
Published: (2025)
A Vision for Multisensory Intelligence: Sensing, Science, and Synergy
by: Liang, Paul Pu
Published: (2026)
by: Liang, Paul Pu
Published: (2026)
MultiMed: Massively Multimodal and Multitask Medical Understanding
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023)
by: Mo, Shentong, et al.
Published: (2023)
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
by: Sun, Shilin, et al.
Published: (2024)
by: Sun, Shilin, et al.
Published: (2024)
OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models
by: Xue, Yida, et al.
Published: (2026)
by: Xue, Yida, et al.
Published: (2026)
Scaling Spatial Intelligence with Multimodal Foundation Models
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
To Forget or Not? Towards Practical Knowledge Unlearning for Large Language Models
by: Tian, Bozhong, et al.
Published: (2024)
by: Tian, Bozhong, et al.
Published: (2024)
LookAhead Tuning: Safer Language Models via Partial Answer Previews
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective
by: Zhu, Xiangru, et al.
Published: (2024)
by: Zhu, Xiangru, et al.
Published: (2024)
WorDepth: Variational Language Prior for Monocular Depth Estimation
by: Zeng, Ziyao, et al.
Published: (2024)
by: Zeng, Ziyao, et al.
Published: (2024)
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
by: Li, Baiqi, et al.
Published: (2024)
by: Li, Baiqi, et al.
Published: (2024)
State Space Model for New-Generation Network Alternative to Transformers: A Survey
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
by: Deng, Ailin, et al.
Published: (2024)
by: Deng, Ailin, et al.
Published: (2024)
NVLM: Open Frontier-Class Multimodal LLMs
by: Dai, Wenliang, et al.
Published: (2024)
by: Dai, Wenliang, et al.
Published: (2024)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
by: Xing, Zhen, et al.
Published: (2024)
by: Xing, Zhen, et al.
Published: (2024)
SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation
by: Nag, Sayan, et al.
Published: (2024)
by: Nag, Sayan, et al.
Published: (2024)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
by: Zhang, Yi-Kai, et al.
Published: (2024)
by: Zhang, Yi-Kai, et al.
Published: (2024)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
Rethinking Prompting Strategies for Multi-Label Recognition with Partial Annotations
by: Rawlekar, Samyak, et al.
Published: (2024)
by: Rawlekar, Samyak, et al.
Published: (2024)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models
by: Hao, Haoran, et al.
Published: (2024)
by: Hao, Haoran, et al.
Published: (2024)
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning
by: Zhong, Yiwu, et al.
Published: (2024)
by: Zhong, Yiwu, et al.
Published: (2024)
Heracles: A Hybrid SSM-Transformer Model for High-Resolution Image and Time-Series Analysis
by: Patro, Badri N., et al.
Published: (2024)
by: Patro, Badri N., et al.
Published: (2024)
ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
by: Duan, Chengqi, et al.
Published: (2025)
by: Duan, Chengqi, et al.
Published: (2025)
OmniGAIA: Towards Native Omni-Modal AI Agents
by: Li, Xiaoxi, et al.
Published: (2026)
by: Li, Xiaoxi, et al.
Published: (2026)
Contrastive Visual Data Augmentation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
Multimodal Long Video Modeling Based on Temporal Dynamic Context
by: Hao, Haoran, et al.
Published: (2025)
by: Hao, Haoran, et al.
Published: (2025)
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
by: Zhang, Renrui, et al.
Published: (2023)
by: Zhang, Renrui, et al.
Published: (2023)
Aligning Agentic World Models via Knowledgeable Experience Learning
by: Ren, Baochang, et al.
Published: (2026)
by: Ren, Baochang, et al.
Published: (2026)
Machine Unlearning in Hyperbolic vs. Euclidean Multimodal Contrastive Learning: Adapting Alignment Calibration to MERU
by: Vidal, Àlex Pujol, et al.
Published: (2025)
by: Vidal, Àlex Pujol, et al.
Published: (2025)
Finding Distributed Object-Centric Properties in Self-Supervised Transformers
by: Rawlekar, Samyak, et al.
Published: (2026)
by: Rawlekar, Samyak, et al.
Published: (2026)
Towards Understanding Camera Motions in Any Video
by: Lin, Zhiqiu, et al.
Published: (2025)
by: Lin, Zhiqiu, et al.
Published: (2025)
OneLLM: One Framework to Align All Modalities with Language
by: Han, Jiaming, et al.
Published: (2023)
by: Han, Jiaming, et al.
Published: (2023)
Improving Prediction Performance and Model Interpretability through Attention Mechanisms from Basic and Applied Research Perspectives
by: Kitada, Shunsuke
Published: (2023)
by: Kitada, Shunsuke
Published: (2023)
Similar Items
-
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024) -
HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and Reasoning
by: Zhou, Chuhao, et al.
Published: (2025) -
A Vision for Multisensory Intelligence: Sensing, Science, and Synergy
by: Liang, Paul Pu
Published: (2026) -
MultiMed: Massively Multimodal and Multitask Medical Understanding
by: Mo, Shentong, et al.
Published: (2024) -
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023)