Guardado en:
| Autores principales: | Pang, Haozhou, Ding, Tianwei, He, Lanshan, Gan, Qi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2503.09645 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
por: Pang, Haozhou, et al.
Publicado: (2024)
por: Pang, Haozhou, et al.
Publicado: (2024)
TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography
por: Dai, Yuqin, et al.
Publicado: (2025)
por: Dai, Yuqin, et al.
Publicado: (2025)
Lodge++: High-quality and Long Dance Generation with Vivid Choreography Patterns
por: Li, Ronghui, et al.
Publicado: (2024)
por: Li, Ronghui, et al.
Publicado: (2024)
A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
por: Zaouali, Mahmoud Chick, et al.
Publicado: (2025)
por: Zaouali, Mahmoud Chick, et al.
Publicado: (2025)
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
por: Wang, Yiping, et al.
Publicado: (2024)
por: Wang, Yiping, et al.
Publicado: (2024)
Towards Understanding Graphical Perception in Large Multimodal Models
por: Zhang, Kai, et al.
Publicado: (2025)
por: Zhang, Kai, et al.
Publicado: (2025)
SMooGPT: Stylized Motion Generation using Large Language Models
por: Zhong, Lei, et al.
Publicado: (2025)
por: Zhong, Lei, et al.
Publicado: (2025)
SplatFont3D: Structure-Aware Text-to-3D Artistic Font Generation with Part-Level Style Control
por: Gan, Ji, et al.
Publicado: (2025)
por: Gan, Ji, et al.
Publicado: (2025)
Neural Cone Radiosity for Interactive Global Illumination with Glossy Materials
por: Ren, Jierui, et al.
Publicado: (2025)
por: Ren, Jierui, et al.
Publicado: (2025)
MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
por: Fang, Shuangkang, et al.
Publicado: (2025)
por: Fang, Shuangkang, et al.
Publicado: (2025)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
por: Jing, Liqiang, et al.
Publicado: (2025)
por: Jing, Liqiang, et al.
Publicado: (2025)
Is this chart lying to me? Automating the detection of misleading visualizations
por: Tonglet, Jonathan, et al.
Publicado: (2025)
por: Tonglet, Jonathan, et al.
Publicado: (2025)
TexGS-VolVis: Expressive Scene Editing for Volume Visualization via Textured Gaussian Splatting
por: Tang, Kaiyuan, et al.
Publicado: (2025)
por: Tang, Kaiyuan, et al.
Publicado: (2025)
FlairGPT: Repurposing LLMs for Interior Designs
por: Littlefair, Gabrielle, et al.
Publicado: (2025)
por: Littlefair, Gabrielle, et al.
Publicado: (2025)
Co-Layout: LLM-driven Co-optimization for Interior Layout
por: Xiang, Chucheng, et al.
Publicado: (2025)
por: Xiang, Chucheng, et al.
Publicado: (2025)
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
por: Gal, Rinon, et al.
Publicado: (2024)
por: Gal, Rinon, et al.
Publicado: (2024)
T$^3$-S2S: Training-free Triplet Tuning for Sketch to Scene Synthesis in Controllable Concept Art Generation
por: Sun, Zhenhong, et al.
Publicado: (2024)
por: Sun, Zhenhong, et al.
Publicado: (2024)
CAP: Evaluation of Persuasive and Creative Image Generation
por: Aghazadeh, Aysan, et al.
Publicado: (2024)
por: Aghazadeh, Aysan, et al.
Publicado: (2024)
TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
por: Guan, Jiazhi, et al.
Publicado: (2024)
por: Guan, Jiazhi, et al.
Publicado: (2024)
OccluGaussian: Occlusion-Aware Gaussian Splatting for Large Scene Reconstruction and Rendering
por: Liu, Shiyong, et al.
Publicado: (2025)
por: Liu, Shiyong, et al.
Publicado: (2025)
Human-Aware 3D Scene Generation with Spatially-constrained Diffusion Models
por: Hong, Xiaolin, et al.
Publicado: (2024)
por: Hong, Xiaolin, et al.
Publicado: (2024)
Grounding Language in Multi-Perspective Referential Communication
por: Tang, Zineng, et al.
Publicado: (2024)
por: Tang, Zineng, et al.
Publicado: (2024)
ORACLE: Orchestrate NPC Daily Activities using Contrastive Learning with Transformer-CVAE
por: Hong, Seong-Eun, et al.
Publicado: (2026)
por: Hong, Seong-Eun, et al.
Publicado: (2026)
VLMaterial: Procedural Material Generation with Large Vision-Language Models
por: Li, Beichen, et al.
Publicado: (2025)
por: Li, Beichen, et al.
Publicado: (2025)
PairingNet: A Learning-based Pair-searching and -matching Network for Image Fragments
por: Zhou, Rixin, et al.
Publicado: (2023)
por: Zhou, Rixin, et al.
Publicado: (2023)
A Simplified Positional Cell Type Visualization using Spatially Aggregated Clusters
por: Mason, Lee, et al.
Publicado: (2024)
por: Mason, Lee, et al.
Publicado: (2024)
Inverse Rendering using Multi-Bounce Path Tracing and Reservoir Sampling
por: Dai, Yuxin, et al.
Publicado: (2024)
por: Dai, Yuxin, et al.
Publicado: (2024)
EAG-PT: Emission-Aware Gaussians and Path Tracing for Diffuse Indoor Scene Reconstruction and Editing
por: Yang, Xijie, et al.
Publicado: (2026)
por: Yang, Xijie, et al.
Publicado: (2026)
Image Generation Models: A Technical History
por: Shirvani, Rouzbeh
Publicado: (2026)
por: Shirvani, Rouzbeh
Publicado: (2026)
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models
por: Wu, Ronghuan, et al.
Publicado: (2024)
por: Wu, Ronghuan, et al.
Publicado: (2024)
PALP: Prompt Aligned Personalization of Text-to-Image Models
por: Arar, Moab, et al.
Publicado: (2024)
por: Arar, Moab, et al.
Publicado: (2024)
FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On
por: Karras, Johanna, et al.
Publicado: (2026)
por: Karras, Johanna, et al.
Publicado: (2026)
Real-Time Position-Aware View Synthesis from Single-View Input
por: Gond, Manu, et al.
Publicado: (2024)
por: Gond, Manu, et al.
Publicado: (2024)
Taking Language Embedded 3D Gaussian Splatting into the Wild
por: Wang, Yuze, et al.
Publicado: (2025)
por: Wang, Yuze, et al.
Publicado: (2025)
LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPS
por: Li, Wanhua, et al.
Publicado: (2025)
por: Li, Wanhua, et al.
Publicado: (2025)
FontCLIP: A Semantic Typography Visual-Language Model for Multilingual Font Applications
por: Tatsukawa, Yuki, et al.
Publicado: (2024)
por: Tatsukawa, Yuki, et al.
Publicado: (2024)
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
por: Wang, Yuze, et al.
Publicado: (2025)
por: Wang, Yuze, et al.
Publicado: (2025)
CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models
por: Pyatov, Vladislav, et al.
Publicado: (2026)
por: Pyatov, Vladislav, et al.
Publicado: (2026)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
por: Mukhopadhyay, Srija, et al.
Publicado: (2024)
por: Mukhopadhyay, Srija, et al.
Publicado: (2024)
Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation
por: He, Lanshan, et al.
Publicado: (2026)
por: He, Lanshan, et al.
Publicado: (2026)
Ejemplares similares
-
LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
por: Pang, Haozhou, et al.
Publicado: (2024) -
TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography
por: Dai, Yuqin, et al.
Publicado: (2025) -
Lodge++: High-quality and Long Dance Generation with Vivid Choreography Patterns
por: Li, Ronghui, et al.
Publicado: (2024) -
A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
por: Zaouali, Mahmoud Chick, et al.
Publicado: (2025) -
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
por: Wang, Yiping, et al.
Publicado: (2024)