GPIC: A Giant Permissive Image Corpus for Visual Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chandrasegaran, Keshigeyan, Sargent, Kyle, Agarwal, Suchir, Jang, Michael, Poli, Michael, Niebles, Juan Carlos, Johnson, Justin, Wu, Jiajun, Fei-Fei, Li |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization
by: Sargent, Kyle, et al.
Published: (2025)
by: Sargent, Kyle, et al.
Published: (2025)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025)
by: Ye, Jinhui, et al.
Published: (2025)
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
by: Baade, Alan, et al.
Published: (2026)
by: Baade, Alan, et al.
Published: (2026)
HourVideo: 1-Hour Video-Language Understanding
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024)
VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
by: Sargent, Kyle, et al.
Published: (2025)
by: Sargent, Kyle, et al.
Published: (2025)
A Survey on Generative Modeling with Limited Data, Few Shots, and Zero Shot
by: Abdollahzadeh, Milad, et al.
Published: (2023)
by: Abdollahzadeh, Milad, et al.
Published: (2023)
Exploring Diffusion Transformer Designs via Grafting
by: Chandrasegaran, Keshigeyan, et al.
Published: (2025)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2025)
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
MindCube: Spatial Mental Modeling from Limited Views
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
Model Inversion Robustness: Can Transfer Learning Help?
by: Ho, Sy-Tuyen, et al.
Published: (2024)
by: Ho, Sy-Tuyen, et al.
Published: (2024)
Understanding Complexity in VideoQA via Visual Program Generation
by: Eyzaguirre, Cristobal, et al.
Published: (2025)
by: Eyzaguirre, Cristobal, et al.
Published: (2025)
ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image
by: Sargent, Kyle, et al.
Published: (2023)
by: Sargent, Kyle, et al.
Published: (2023)
PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models
by: Fei, Fan, et al.
Published: (2025)
by: Fei, Fan, et al.
Published: (2025)
Streaming Detection of Queried Event Start
by: Eyzaguirre, Cristobal, et al.
Published: (2024)
by: Eyzaguirre, Cristobal, et al.
Published: (2024)
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
by: Dharmarajan, Karthik, et al.
Published: (2025)
by: Dharmarajan, Karthik, et al.
Published: (2025)
WorldScore: A Unified Evaluation Benchmark for World Generation
by: Duan, Haoyi, et al.
Published: (2025)
by: Duan, Haoyi, et al.
Published: (2025)
View-Invariant Policy Learning via Zero-Shot Novel View Synthesis
by: Tian, Stephen, et al.
Published: (2024)
by: Tian, Stephen, et al.
Published: (2024)
ViUniT: Visual Unit Tests for More Robust Visual Programming
by: Panagopoulou, Artemis, et al.
Published: (2024)
by: Panagopoulou, Artemis, et al.
Published: (2024)
AdaVid: Adaptive Video-Language Pretraining
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
SurfPhase: 3D Interfacial Dynamics in Two-Phase Flows from Sparse Videos
by: Gao, Yue, et al.
Published: (2026)
by: Gao, Yue, et al.
Published: (2026)
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
by: Zhang, Pingyue, et al.
Published: (2026)
by: Zhang, Pingyue, et al.
Published: (2026)
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
by: Tang, Yihe, et al.
Published: (2025)
by: Tang, Yihe, et al.
Published: (2025)
WonderJourney: Going from Anywhere to Everywhere
by: Yu, Hong-Xing, et al.
Published: (2023)
by: Yu, Hong-Xing, et al.
Published: (2023)
UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
Unifying Specialized Visual Encoders for Video Language Models
by: Chung, Jihoon, et al.
Published: (2025)
by: Chung, Jihoon, et al.
Published: (2025)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
by: Kendre, Shrikant, et al.
Published: (2025)
by: Kendre, Shrikant, et al.
Published: (2025)
Product of Experts for Visual Generation
by: Zhang, Yunzhi, et al.
Published: (2025)
by: Zhang, Yunzhi, et al.
Published: (2025)
VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
by: Hou, Haowen, et al.
Published: (2024)
by: Hou, Haowen, et al.
Published: (2024)
Future Optical Flow Prediction Improves Robot Control & Video Generation
by: Ranasinghe, Kanchana, et al.
Published: (2026)
by: Ranasinghe, Kanchana, et al.
Published: (2026)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
by: Liu, Yunong, et al.
Published: (2024)
by: Liu, Yunong, et al.
Published: (2024)
Language-Informed Visual Concept Learning
by: Lee, Sharon, et al.
Published: (2023)
by: Lee, Sharon, et al.
Published: (2023)
Taming Generative Diffusion Prior for Universal Blind Image Restoration
by: Tu, Siwei, et al.
Published: (2024)
by: Tu, Siwei, et al.
Published: (2024)
Few-shot Defect Image Generation based on Consistency Modeling
by: Shi, Qingfeng, et al.
Published: (2024)
by: Shi, Qingfeng, et al.
Published: (2024)
Cloud Adversarial Example Generation for Remote Sensing Image Classification
by: Ma, Fei, et al.
Published: (2024)
by: Ma, Fei, et al.
Published: (2024)
VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
by: Yin, Shaofeng, et al.
Published: (2025)
by: Yin, Shaofeng, et al.
Published: (2025)
Universal Scene Graph Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Cascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment
by: Yang, Lei, et al.
Published: (2026)
by: Yang, Lei, et al.
Published: (2026)
Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection
by: Lin, Kaiqing, et al.
Published: (2024)
by: Lin, Kaiqing, et al.
Published: (2024)
Safe-VAR: Safe Visual Autoregressive Model for Text-to-Image Generative Watermarking
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
M3: High-fidelity Text-to-Image Generation via Multi-Modal, Multi-Agent and Multi-Round Visual Reasoning
by: Yang, Bangji, et al.
Published: (2026)
by: Yang, Bangji, et al.
Published: (2026)
Similar Items
-
Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization
by: Sargent, Kyle, et al.
Published: (2025) -
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025) -
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
by: Baade, Alan, et al.
Published: (2026) -
HourVideo: 1-Hour Video-Language Understanding
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024) -
VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
by: Sargent, Kyle, et al.
Published: (2025)