Don't Judge Before You CLIP: A Unified Approach for Perceptual Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Zalcher, Amit, Wasserman, Navve, Beliy, Roman, Heinimann, Oliver, Irani, Michal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Wisdom of a Crowd of Brains: A Universal Brain Encoder
by: Beliy, Roman, et al.
Published: (2024)
by: Beliy, Roman, et al.
Published: (2024)
Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer
by: Beliy, Roman, et al.
Published: (2025)
by: Beliy, Roman, et al.
Published: (2025)
Brain-IT-VQA: From Brain Signals to Answers
by: Beliy, Roman, et al.
Published: (2026)
by: Beliy, Roman, et al.
Published: (2026)
From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain
by: Golbari, Yuval, et al.
Published: (2026)
by: Golbari, Yuval, et al.
Published: (2026)
ImpMIA: Leveraging Implicit Bias for Membership Inference Attack
by: Golbari, Yuval, et al.
Published: (2025)
by: Golbari, Yuval, et al.
Published: (2025)
KernelFusion: Assumption-Free Blind Super-Resolution via Patch Diffusion
by: Heinimann, Oliver, et al.
Published: (2025)
by: Heinimann, Oliver, et al.
Published: (2025)
BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain
by: Wasserman, Navve, et al.
Published: (2025)
by: Wasserman, Navve, et al.
Published: (2025)
Functional Brain-to-Brain Transformation with No Shared Data
by: Wasserman, Navve, et al.
Published: (2024)
by: Wasserman, Navve, et al.
Published: (2024)
NeIn: Telling What You Don't Want
by: Bui, Nhat-Tan, et al.
Published: (2024)
by: Bui, Nhat-Tan, et al.
Published: (2024)
Now You See Me, Now You Don't: A Unified Framework for Expression Consistent Anonymization in Talking Head Videos
by: Egin, Anil, et al.
Published: (2026)
by: Egin, Anil, et al.
Published: (2026)
Perceptual Inductive Bias Is What You Need Before Contrastive Learning
by: Li, Tianqin, et al.
Published: (2025)
by: Li, Tianqin, et al.
Published: (2025)
Paint by Inpaint: Learning to Add Image Objects by Removing Them First
by: Wasserman, Navve, et al.
Published: (2024)
by: Wasserman, Navve, et al.
Published: (2024)
Visually Dehallucinative Instruction Generation: Know What You Don't Know
by: Cha, Sungguk, et al.
Published: (2024)
by: Cha, Sungguk, et al.
Published: (2024)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024)
by: Li, Senmao, et al.
Published: (2024)
Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation
by: Biswas, Shristi Das, et al.
Published: (2025)
by: Biswas, Shristi Das, et al.
Published: (2025)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
by: Wang, Juan, et al.
Published: (2026)
by: Wang, Juan, et al.
Published: (2026)
When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
by: Tuchinda, Pume, et al.
Published: (2025)
by: Tuchinda, Pume, et al.
Published: (2025)
If At First You Don't Succeed: Test Time Re-ranking for Zero-shot, Cross-domain Retrieval
by: Hudson, Finlay G. C., et al.
Published: (2023)
by: Hudson, Finlay G. C., et al.
Published: (2023)
Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
by: Kuckreja, Kartik, et al.
Published: (2026)
by: Kuckreja, Kartik, et al.
Published: (2026)
You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning
by: Moutakanni, Théo, et al.
Published: (2024)
by: Moutakanni, Théo, et al.
Published: (2024)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
REAL-MM-RAG: A Real-World Multi-Modal Retrieval Benchmark
by: Wasserman, Navve, et al.
Published: (2025)
by: Wasserman, Navve, et al.
Published: (2025)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
by: Chen, Harold Haodong, et al.
Published: (2026)
by: Chen, Harold Haodong, et al.
Published: (2026)
Don't Pause! Every prediction matters in a streaming video
by: Chatterjee, Dibyadip, et al.
Published: (2026)
by: Chatterjee, Dibyadip, et al.
Published: (2026)
Don't Look into the Dark: Latent Codes for Pluralistic Image Inpainting
by: Chen, Haiwei, et al.
Published: (2024)
by: Chen, Haiwei, et al.
Published: (2024)
ProCreate, Don't Reproduce! Propulsive Energy Diffusion for Creative Generation
by: Lu, Jack, et al.
Published: (2024)
by: Lu, Jack, et al.
Published: (2024)
Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
by: Ci, En, et al.
Published: (2025)
by: Ci, En, et al.
Published: (2025)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankers
by: Wasserman, Navve, et al.
Published: (2025)
by: Wasserman, Navve, et al.
Published: (2025)
VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
by: Sargent, Kyle, et al.
Published: (2025)
by: Sargent, Kyle, et al.
Published: (2025)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Don't let the information slip away
by: Li, Taozhe, et al.
Published: (2026)
by: Li, Taozhe, et al.
Published: (2026)
Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs
by: Boroujeni, Sayed Pedram Haeri, et al.
Published: (2026)
by: Boroujeni, Sayed Pedram Haeri, et al.
Published: (2026)
Don't Look at the Camera: Achieving Perceived Eye Contact
by: Gao, Alice, et al.
Published: (2024)
by: Gao, Alice, et al.
Published: (2024)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
by: Zhao, Canyu, et al.
Published: (2025)
by: Zhao, Canyu, et al.
Published: (2025)
Don't Fear Peculiar Activation Functions: EUAF and Beyond
by: Wang, Qianchao, et al.
Published: (2024)
by: Wang, Qianchao, et al.
Published: (2024)
Don't Mind the Gaps: Implicit Neural Representations for Resolution-Agnostic Retinal OCT Analysis
by: Kahrs, Bennet, et al.
Published: (2026)
by: Kahrs, Bennet, et al.
Published: (2026)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
by: Zou, Xin, et al.
Published: (2025)
by: Zou, Xin, et al.
Published: (2025)
Think Before You Diffuse: Infusing Physical Rules into Video Diffusion
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
Similar Items
-
The Wisdom of a Crowd of Brains: A Universal Brain Encoder
by: Beliy, Roman, et al.
Published: (2024) -
Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer
by: Beliy, Roman, et al.
Published: (2025) -
Brain-IT-VQA: From Brain Signals to Answers
by: Beliy, Roman, et al.
Published: (2026) -
From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain
by: Golbari, Yuval, et al.
Published: (2026) -
ImpMIA: Leveraging Implicit Bias for Membership Inference Attack
by: Golbari, Yuval, et al.
Published: (2025)