Enregistré dans:
| Auteurs principaux: | Pang, Haozhou, Ding, Tianwei, He, Lanshan, Gan, Qi |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2503.09645 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
par: Pang, Haozhou, et autres
Publié: (2024)
par: Pang, Haozhou, et autres
Publié: (2024)
TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography
par: Dai, Yuqin, et autres
Publié: (2025)
par: Dai, Yuqin, et autres
Publié: (2025)
Lodge++: High-quality and Long Dance Generation with Vivid Choreography Patterns
par: Li, Ronghui, et autres
Publié: (2024)
par: Li, Ronghui, et autres
Publié: (2024)
A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
par: Zaouali, Mahmoud Chick, et autres
Publié: (2025)
par: Zaouali, Mahmoud Chick, et autres
Publié: (2025)
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
par: Wang, Yiping, et autres
Publié: (2024)
par: Wang, Yiping, et autres
Publié: (2024)
Towards Understanding Graphical Perception in Large Multimodal Models
par: Zhang, Kai, et autres
Publié: (2025)
par: Zhang, Kai, et autres
Publié: (2025)
SMooGPT: Stylized Motion Generation using Large Language Models
par: Zhong, Lei, et autres
Publié: (2025)
par: Zhong, Lei, et autres
Publié: (2025)
SplatFont3D: Structure-Aware Text-to-3D Artistic Font Generation with Part-Level Style Control
par: Gan, Ji, et autres
Publié: (2025)
par: Gan, Ji, et autres
Publié: (2025)
Neural Cone Radiosity for Interactive Global Illumination with Glossy Materials
par: Ren, Jierui, et autres
Publié: (2025)
par: Ren, Jierui, et autres
Publié: (2025)
MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
par: Fang, Shuangkang, et autres
Publié: (2025)
par: Fang, Shuangkang, et autres
Publié: (2025)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
par: Jing, Liqiang, et autres
Publié: (2025)
par: Jing, Liqiang, et autres
Publié: (2025)
Is this chart lying to me? Automating the detection of misleading visualizations
par: Tonglet, Jonathan, et autres
Publié: (2025)
par: Tonglet, Jonathan, et autres
Publié: (2025)
TexGS-VolVis: Expressive Scene Editing for Volume Visualization via Textured Gaussian Splatting
par: Tang, Kaiyuan, et autres
Publié: (2025)
par: Tang, Kaiyuan, et autres
Publié: (2025)
FlairGPT: Repurposing LLMs for Interior Designs
par: Littlefair, Gabrielle, et autres
Publié: (2025)
par: Littlefair, Gabrielle, et autres
Publié: (2025)
Co-Layout: LLM-driven Co-optimization for Interior Layout
par: Xiang, Chucheng, et autres
Publié: (2025)
par: Xiang, Chucheng, et autres
Publié: (2025)
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
par: Gal, Rinon, et autres
Publié: (2024)
par: Gal, Rinon, et autres
Publié: (2024)
T$^3$-S2S: Training-free Triplet Tuning for Sketch to Scene Synthesis in Controllable Concept Art Generation
par: Sun, Zhenhong, et autres
Publié: (2024)
par: Sun, Zhenhong, et autres
Publié: (2024)
CAP: Evaluation of Persuasive and Creative Image Generation
par: Aghazadeh, Aysan, et autres
Publié: (2024)
par: Aghazadeh, Aysan, et autres
Publié: (2024)
TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
par: Guan, Jiazhi, et autres
Publié: (2024)
par: Guan, Jiazhi, et autres
Publié: (2024)
OccluGaussian: Occlusion-Aware Gaussian Splatting for Large Scene Reconstruction and Rendering
par: Liu, Shiyong, et autres
Publié: (2025)
par: Liu, Shiyong, et autres
Publié: (2025)
Human-Aware 3D Scene Generation with Spatially-constrained Diffusion Models
par: Hong, Xiaolin, et autres
Publié: (2024)
par: Hong, Xiaolin, et autres
Publié: (2024)
Grounding Language in Multi-Perspective Referential Communication
par: Tang, Zineng, et autres
Publié: (2024)
par: Tang, Zineng, et autres
Publié: (2024)
ORACLE: Orchestrate NPC Daily Activities using Contrastive Learning with Transformer-CVAE
par: Hong, Seong-Eun, et autres
Publié: (2026)
par: Hong, Seong-Eun, et autres
Publié: (2026)
VLMaterial: Procedural Material Generation with Large Vision-Language Models
par: Li, Beichen, et autres
Publié: (2025)
par: Li, Beichen, et autres
Publié: (2025)
PairingNet: A Learning-based Pair-searching and -matching Network for Image Fragments
par: Zhou, Rixin, et autres
Publié: (2023)
par: Zhou, Rixin, et autres
Publié: (2023)
A Simplified Positional Cell Type Visualization using Spatially Aggregated Clusters
par: Mason, Lee, et autres
Publié: (2024)
par: Mason, Lee, et autres
Publié: (2024)
Inverse Rendering using Multi-Bounce Path Tracing and Reservoir Sampling
par: Dai, Yuxin, et autres
Publié: (2024)
par: Dai, Yuxin, et autres
Publié: (2024)
EAG-PT: Emission-Aware Gaussians and Path Tracing for Diffuse Indoor Scene Reconstruction and Editing
par: Yang, Xijie, et autres
Publié: (2026)
par: Yang, Xijie, et autres
Publié: (2026)
Image Generation Models: A Technical History
par: Shirvani, Rouzbeh
Publié: (2026)
par: Shirvani, Rouzbeh
Publié: (2026)
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models
par: Wu, Ronghuan, et autres
Publié: (2024)
par: Wu, Ronghuan, et autres
Publié: (2024)
PALP: Prompt Aligned Personalization of Text-to-Image Models
par: Arar, Moab, et autres
Publié: (2024)
par: Arar, Moab, et autres
Publié: (2024)
FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On
par: Karras, Johanna, et autres
Publié: (2026)
par: Karras, Johanna, et autres
Publié: (2026)
Real-Time Position-Aware View Synthesis from Single-View Input
par: Gond, Manu, et autres
Publié: (2024)
par: Gond, Manu, et autres
Publié: (2024)
Taking Language Embedded 3D Gaussian Splatting into the Wild
par: Wang, Yuze, et autres
Publié: (2025)
par: Wang, Yuze, et autres
Publié: (2025)
LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPS
par: Li, Wanhua, et autres
Publié: (2025)
par: Li, Wanhua, et autres
Publié: (2025)
FontCLIP: A Semantic Typography Visual-Language Model for Multilingual Font Applications
par: Tatsukawa, Yuki, et autres
Publié: (2024)
par: Tatsukawa, Yuki, et autres
Publié: (2024)
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
par: Wang, Yuze, et autres
Publié: (2025)
par: Wang, Yuze, et autres
Publié: (2025)
CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models
par: Pyatov, Vladislav, et autres
Publié: (2026)
par: Pyatov, Vladislav, et autres
Publié: (2026)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
par: Mukhopadhyay, Srija, et autres
Publié: (2024)
par: Mukhopadhyay, Srija, et autres
Publié: (2024)
Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation
par: He, Lanshan, et autres
Publié: (2026)
par: He, Lanshan, et autres
Publié: (2026)
Documents similaires
-
LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
par: Pang, Haozhou, et autres
Publié: (2024) -
TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography
par: Dai, Yuqin, et autres
Publié: (2025) -
Lodge++: High-quality and Long Dance Generation with Vivid Choreography Patterns
par: Li, Ronghui, et autres
Publié: (2024) -
A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
par: Zaouali, Mahmoud Chick, et autres
Publié: (2025) -
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
par: Wang, Yiping, et autres
Publié: (2024)