HOIN: High-Order Implicit Neural Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yang, Wu, Ruituo, Liu, Yipeng, Zhu, Ce |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Long-tailed Medical Diagnosis with Relation-aware Representation Learning and Iterative Classifier Calibration
von: Pan, Li, et al.
Veröffentlicht: (2025)
von: Pan, Li, et al.
Veröffentlicht: (2025)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Improving Visual Representation Alignment Generation with GRPO
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
von: Li, Liupeng, et al.
Veröffentlicht: (2026)
von: Li, Liupeng, et al.
Veröffentlicht: (2026)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
von: Zhou, Shengli, et al.
Veröffentlicht: (2026)
von: Zhou, Shengli, et al.
Veröffentlicht: (2026)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
von: Yang, Chenglin, et al.
Veröffentlicht: (2023)
von: Yang, Chenglin, et al.
Veröffentlicht: (2023)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
von: Chen, Weifeng, et al.
Veröffentlicht: (2023)
von: Chen, Weifeng, et al.
Veröffentlicht: (2023)
HER2 Expression Prediction with Flexible Multi-Modal Inputs via Dynamic Bidirectional Reconstruction
von: Qin, Jie, et al.
Veröffentlicht: (2025)
von: Qin, Jie, et al.
Veröffentlicht: (2025)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
von: Xiao, Xinyu, et al.
Veröffentlicht: (2026)
von: Xiao, Xinyu, et al.
Veröffentlicht: (2026)
ClassWise-CRF: Category-Specific Fusion for Enhanced Semantic Segmentation of Remote Sensing Imagery
von: Zhu, Qinfeng, et al.
Veröffentlicht: (2025)
von: Zhu, Qinfeng, et al.
Veröffentlicht: (2025)
STIV: Scalable Text and Image Conditioned Video Generation
von: Lin, Zongyu, et al.
Veröffentlicht: (2024)
von: Lin, Zongyu, et al.
Veröffentlicht: (2024)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
von: Lin, Yuanze, et al.
Veröffentlicht: (2025)
von: Lin, Yuanze, et al.
Veröffentlicht: (2025)
Evaluating the Impact of Point Cloud Colorization on Semantic Segmentation Accuracy
von: Zhu, Qinfeng, et al.
Veröffentlicht: (2024)
von: Zhu, Qinfeng, et al.
Veröffentlicht: (2024)
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
von: Wu, Xiaoping, et al.
Veröffentlicht: (2024)
von: Wu, Xiaoping, et al.
Veröffentlicht: (2024)
Vlogger: Make Your Dream A Vlog
von: Zhuang, Shaobin, et al.
Veröffentlicht: (2024)
von: Zhuang, Shaobin, et al.
Veröffentlicht: (2024)
Size Matters: Reconstructing Real-Scale 3D Models from Monocular Images for Food Portion Estimation
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
von: Zhang, Yabin, et al.
Veröffentlicht: (2024)
von: Zhang, Yabin, et al.
Veröffentlicht: (2024)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025)
HeCoFuse: Cross-Modal Complementary V2X Cooperative Perception with Heterogeneous Sensors
von: Wei, Chuheng, et al.
Veröffentlicht: (2025)
von: Wei, Chuheng, et al.
Veröffentlicht: (2025)
How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic Segmentation
von: Pan, Yining, et al.
Veröffentlicht: (2025)
von: Pan, Yining, et al.
Veröffentlicht: (2025)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
von: Li, Quanhao, et al.
Veröffentlicht: (2025)
von: Li, Quanhao, et al.
Veröffentlicht: (2025)
SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation
von: Li, Xuewei, et al.
Veröffentlicht: (2023)
von: Li, Xuewei, et al.
Veröffentlicht: (2023)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
von: Huang, Po-Hsuan, et al.
Veröffentlicht: (2024)
von: Huang, Po-Hsuan, et al.
Veröffentlicht: (2024)
ReconBoost: Boosting Can Achieve Modality Reconcilement
von: Hua, Cong, et al.
Veröffentlicht: (2024)
von: Hua, Cong, et al.
Veröffentlicht: (2024)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
von: Li, Quanhao, et al.
Veröffentlicht: (2026)
von: Li, Quanhao, et al.
Veröffentlicht: (2026)
Explore the Limits of Omni-modal Pretraining at Scale
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
von: Liu, Sheng, et al.
Veröffentlicht: (2024)
von: Liu, Sheng, et al.
Veröffentlicht: (2024)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
von: Chen, Baiyu, et al.
Veröffentlicht: (2025)
von: Chen, Baiyu, et al.
Veröffentlicht: (2025)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
von: Tang, Shixiang, et al.
Veröffentlicht: (2025)
von: Tang, Shixiang, et al.
Veröffentlicht: (2025)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
von: Wang, Zhouxia, et al.
Veröffentlicht: (2023)
von: Wang, Zhouxia, et al.
Veröffentlicht: (2023)
Pay Less Attention to Deceptive Artifacts: Robust Detection of Compressed Deepfakes on Online Social Networks
von: Li, Manyi, et al.
Veröffentlicht: (2025)
von: Li, Manyi, et al.
Veröffentlicht: (2025)
LayerT2V: A Unified Multi-Layer Video Generation Framework
von: Li, Guangzhao, et al.
Veröffentlicht: (2025)
von: Li, Guangzhao, et al.
Veröffentlicht: (2025)
A Systematic Review on Long-Tailed Learning
von: Zhang, Chongsheng, et al.
Veröffentlicht: (2024)
von: Zhang, Chongsheng, et al.
Veröffentlicht: (2024)
UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
von: Li, Junzhe, et al.
Veröffentlicht: (2025)
von: Li, Junzhe, et al.
Veröffentlicht: (2025)
GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models
von: Zhang, Zewei, et al.
Veröffentlicht: (2024)
von: Zhang, Zewei, et al.
Veröffentlicht: (2024)
FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Long-tailed Medical Diagnosis with Relation-aware Representation Learning and Iterative Classifier Calibration
von: Pan, Li, et al.
Veröffentlicht: (2025) -
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
von: Chen, Yang, et al.
Veröffentlicht: (2024) -
Improving Visual Representation Alignment Generation with GRPO
von: Mo, Shentong, et al.
Veröffentlicht: (2026) -
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
von: Li, Liupeng, et al.
Veröffentlicht: (2026) -
PlanLLM: Video Procedure Planning with Refinable Large Language Models
von: Yang, Dejie, et al.
Veröffentlicht: (2024)