Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Cheng, Shihan, Kulkarni, Nilesh, Hyde, David, Smirnov, Dmitriy |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
par: Kazemi, Amir, et autres
Publié: (2024)
par: Kazemi, Amir, et autres
Publié: (2024)
A Hybrid Multimodal Deep Learning Framework for Intelligent Fashion Recommendation
par: Kalashi, Kamand, et autres
Publié: (2025)
par: Kalashi, Kamand, et autres
Publié: (2025)
PhysMorph-GS: Render-Guided Volumetric Morphing with Differentiable Physics
par: Song, Chang-Yong, et autres
Publié: (2025)
par: Song, Chang-Yong, et autres
Publié: (2025)
TexTile: A Differentiable Metric for Texture Tileability
par: Rodriguez-Pardo, Carlos, et autres
Publié: (2024)
par: Rodriguez-Pardo, Carlos, et autres
Publié: (2024)
Deep Spectral Meshes: Multi-Frequency Facial Mesh Processing with Graph Neural Networks
par: Kosk, Robert, et autres
Publié: (2024)
par: Kosk, Robert, et autres
Publié: (2024)
Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal Prediction
par: Artru, Noé, et autres
Publié: (2026)
par: Artru, Noé, et autres
Publié: (2026)
ROI-GS: Interest-based Local Quality 3D Gaussian Splatting
par: Bui, Quoc-Anh, et autres
Publié: (2025)
par: Bui, Quoc-Anh, et autres
Publié: (2025)
ROI-NeRFs: Hi-Fi Visualization of Objects of Interest within a Scene by NeRFs Composition
par: Bui, Quoc-Anh, et autres
Publié: (2025)
par: Bui, Quoc-Anh, et autres
Publié: (2025)
SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
par: Gondhalekar, Chinmay, et autres
Publié: (2025)
par: Gondhalekar, Chinmay, et autres
Publié: (2025)
Light Field Display Point Rendering
par: Gavane, Ajinkya, et autres
Publié: (2025)
par: Gavane, Ajinkya, et autres
Publié: (2025)
MatDecompSDF: High-Fidelity 3D Shape and PBR Material Decomposition from Multi-View Images
par: Wang, Chengyu, et autres
Publié: (2025)
par: Wang, Chengyu, et autres
Publié: (2025)
Transforming faces into video stories -- VideoFace2.0
par: Brkljač, Branko, et autres
Publié: (2025)
par: Brkljač, Branko, et autres
Publié: (2025)
Sequential PatchCore: Anomaly Detection for Surface Inspection using Synthetic Impurities
par: Mao, Runzhou, et autres
Publié: (2025)
par: Mao, Runzhou, et autres
Publié: (2025)
X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering
par: Huang, Zhitong, et autres
Publié: (2025)
par: Huang, Zhitong, et autres
Publié: (2025)
FLoD: Integrating Flexible Level of Detail into 3D Gaussian Splatting for Customizable Rendering
par: Seo, Yunji, et autres
Publié: (2024)
par: Seo, Yunji, et autres
Publié: (2024)
Structured Basis Function Networks: Loss-Centric Multi-Hypothesis Ensembles with Controllable Diversity
par: Dominguez, Alejandro Rodriguez, et autres
Publié: (2025)
par: Dominguez, Alejandro Rodriguez, et autres
Publié: (2025)
Lightweight MRI-Based Automated Segmentation of Pancreatic Cancer with Auto3DSeg
par: Jha, Keshav, et autres
Publié: (2025)
par: Jha, Keshav, et autres
Publié: (2025)
A New Perspective on Drawing Venn Diagrams for Data Visualization
par: Csanády, Bálint
Publié: (2026)
par: Csanády, Bálint
Publié: (2026)
Inverse 3D Microscopy Rendering for Cell Shape Inference with Active Mesh
par: Ichbiah, Sacha, et autres
Publié: (2023)
par: Ichbiah, Sacha, et autres
Publié: (2023)
APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music
par: Husain, Jaavid Aktar, et autres
Publié: (2026)
par: Husain, Jaavid Aktar, et autres
Publié: (2026)
Real-World En Call Center Transcripts Dataset with PII Redaction
par: Dao, Ha, et autres
Publié: (2025)
par: Dao, Ha, et autres
Publié: (2025)
Advancing Annotat3D with Harpia: A CUDA-Accelerated Library For Large-Scale Volumetric Data Segmentation
par: de Araujo, Camila Machado, et autres
Publié: (2025)
par: de Araujo, Camila Machado, et autres
Publié: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
par: Siddiqui, Yousuf Ahmed, et autres
Publié: (2025)
par: Siddiqui, Yousuf Ahmed, et autres
Publié: (2025)
Advancing Brain Tumor Segmentation via Attention-based 3D U-Net Architecture and Digital Image Processing
par: Gad, Eyad, et autres
Publié: (2025)
par: Gad, Eyad, et autres
Publié: (2025)
How Will It Drape Like? Capturing Fabric Mechanics from Depth Images
par: Rodriguez-Pardo, Carlos, et autres
Publié: (2023)
par: Rodriguez-Pardo, Carlos, et autres
Publié: (2023)
Air Quality Prediction Using LOESS-ARIMA and Multi-Scale CNN-BiLSTM with Residual-Gated Attention
par: Pahari, Soham, et autres
Publié: (2025)
par: Pahari, Soham, et autres
Publié: (2025)
LVCD: Reference-based Lineart Video Colorization with Diffusion Models
par: Huang, Zhitong, et autres
Publié: (2024)
par: Huang, Zhitong, et autres
Publié: (2024)
AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data
par: Kushal, Koushik Ahmed, et autres
Publié: (2025)
par: Kushal, Koushik Ahmed, et autres
Publié: (2025)
HOSC: A Periodic Activation with Saturation Control for High-Fidelity Implicit Neural Representations
par: Wlodarczyk, Michal Jan, et autres
Publié: (2026)
par: Wlodarczyk, Michal Jan, et autres
Publié: (2026)
VocSim: A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio
par: Basha, Maris, et autres
Publié: (2025)
par: Basha, Maris, et autres
Publié: (2025)
Dense Video Understanding with Gated Residual Tokenization
par: Zhang, Haichao, et autres
Publié: (2025)
par: Zhang, Haichao, et autres
Publié: (2025)
Deep Interest Mining for Intent-Enriched Semantic IDs in Multimodal Generative Recommendation
par: Zeng, Yangchen, et autres
Publié: (2026)
par: Zeng, Yangchen, et autres
Publié: (2026)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
par: Patel, Urjitkumar, et autres
Publié: (2025)
par: Patel, Urjitkumar, et autres
Publié: (2025)
LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
par: Zhang, Haichao, et autres
Publié: (2025)
par: Zhang, Haichao, et autres
Publié: (2025)
Reconstructing Curves from Sparse Samples on Riemannian Manifolds
par: Marin, Diana, et autres
Publié: (2024)
par: Marin, Diana, et autres
Publié: (2024)
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
par: Lentsch, Ted, et autres
Publié: (2026)
par: Lentsch, Ted, et autres
Publié: (2026)
UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes
par: Lentsch, Ted, et autres
Publié: (2024)
par: Lentsch, Ted, et autres
Publié: (2024)
Decoder Generates Manufacturable Structures: A Framework for 3D-Printable Object Synthesis
par: Kumar, Abhishek
Publié: (2026)
par: Kumar, Abhishek
Publié: (2026)
Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling
par: Beskorovainyi, Vladimir
Publié: (2026)
par: Beskorovainyi, Vladimir
Publié: (2026)
SShaDe: scalable shape deformation via local representations
par: Maggioli, Filippo, et autres
Publié: (2024)
par: Maggioli, Filippo, et autres
Publié: (2024)
Documents similaires
-
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
par: Kazemi, Amir, et autres
Publié: (2024) -
A Hybrid Multimodal Deep Learning Framework for Intelligent Fashion Recommendation
par: Kalashi, Kamand, et autres
Publié: (2025) -
PhysMorph-GS: Render-Guided Volumetric Morphing with Differentiable Physics
par: Song, Chang-Yong, et autres
Publié: (2025) -
TexTile: A Differentiable Metric for Texture Tileability
par: Rodriguez-Pardo, Carlos, et autres
Publié: (2024) -
Deep Spectral Meshes: Multi-Frequency Facial Mesh Processing with Graph Neural Networks
par: Kosk, Robert, et autres
Publié: (2024)