HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Shengkai, Jiao, Nianhong, Li, Tian, Yang, Chaojie, Xue, Chenhui, Niu, Boya, Gao, Jun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
por: Viveiros, André G., et al.
Publicado: (2025)
por: Viveiros, André G., et al.
Publicado: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
Depth Priors in Removal Neural Radiance Fields
por: Guo, Zhihao, et al.
Publicado: (2024)
por: Guo, Zhihao, et al.
Publicado: (2024)
A Cost-Effective Eye-Tracker for Early Detection of Mild Cognitive Impairment
por: Greco, Danilo, et al.
Publicado: (2024)
por: Greco, Danilo, et al.
Publicado: (2024)
Few-Class Arena: A Benchmark for Efficient Selection of Vision Models and Dataset Difficulty Measurement
por: Cao, Bryan Bo, et al.
Publicado: (2024)
por: Cao, Bryan Bo, et al.
Publicado: (2024)
SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval
por: Wu, Ren-Di, et al.
Publicado: (2025)
por: Wu, Ren-Di, et al.
Publicado: (2025)
Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
por: Kalušev, Vladimir, et al.
Publicado: (2026)
por: Kalušev, Vladimir, et al.
Publicado: (2026)
StatsMerging: Statistics-Guided Model Merging via Task-Specific Teacher Distillation
por: Merugu, Ranjith, et al.
Publicado: (2025)
por: Merugu, Ranjith, et al.
Publicado: (2025)
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases
por: Bell-Navas, Andrés, et al.
Publicado: (2025)
por: Bell-Navas, Andrés, et al.
Publicado: (2025)
Application of deep learning approaches for medieval historical documents transcription
por: Voloshchuk, Maksym, et al.
Publicado: (2025)
por: Voloshchuk, Maksym, et al.
Publicado: (2025)
Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
por: Kang, Xueyang, et al.
Publicado: (2026)
por: Kang, Xueyang, et al.
Publicado: (2026)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
por: Kalkanli, Beyza, et al.
Publicado: (2026)
por: Kalkanli, Beyza, et al.
Publicado: (2026)
Learning Sign Language Representation using CNN LSTM, 3DCNN, CNN RNN LSTM and CCN TD
por: Louison, Nikita, et al.
Publicado: (2024)
por: Louison, Nikita, et al.
Publicado: (2024)
Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal Prediction
por: Artru, Noé, et al.
Publicado: (2026)
por: Artru, Noé, et al.
Publicado: (2026)
Leveraging large multimodal models for audio-video deepfake detection: a pilot study
por: Cao, Songjun, et al.
Publicado: (2026)
por: Cao, Songjun, et al.
Publicado: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
por: Dang, Kieu, et al.
Publicado: (2025)
por: Dang, Kieu, et al.
Publicado: (2025)
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
por: Xi, Wang, et al.
Publicado: (2025)
por: Xi, Wang, et al.
Publicado: (2025)
Transfer learning with generative models for object detection on limited datasets
por: Paiano, Matteo, et al.
Publicado: (2024)
por: Paiano, Matteo, et al.
Publicado: (2024)
Advancing Brain Tumor Segmentation via Attention-based 3D U-Net Architecture and Digital Image Processing
por: Gad, Eyad, et al.
Publicado: (2025)
por: Gad, Eyad, et al.
Publicado: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
por: Pather, Kaviraj, et al.
Publicado: (2025)
por: Pather, Kaviraj, et al.
Publicado: (2025)
Archival Faces: Detection of Faces in Digitized Historical Documents
por: Vaško, Marek, et al.
Publicado: (2025)
por: Vaško, Marek, et al.
Publicado: (2025)
LLM-supported document separation for printed reviews from zbMATH Open
por: Pluzhnikov, Ivan, et al.
Publicado: (2026)
por: Pluzhnikov, Ivan, et al.
Publicado: (2026)
Transforming faces into video stories -- VideoFace2.0
por: Brkljač, Branko, et al.
Publicado: (2025)
por: Brkljač, Branko, et al.
Publicado: (2025)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
por: Berthier, Louis, et al.
Publicado: (2025)
por: Berthier, Louis, et al.
Publicado: (2025)
Identifying Autism-Related Neurobiomarkers Using Hybrid Deep Learning Models
por: Chen, Ashley
Publicado: (2025)
por: Chen, Ashley
Publicado: (2025)
Evaluating Prompting Strategies for Chart Question Answering with Large Language Models
por: Naikar, Ruthuparna, et al.
Publicado: (2026)
por: Naikar, Ruthuparna, et al.
Publicado: (2026)
Enhancing OCR for Sino-Vietnamese Language Processing via Fine-tuned PaddleOCRv5
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
por: Patel, Urjitkumar, et al.
Publicado: (2025)
por: Patel, Urjitkumar, et al.
Publicado: (2025)
ROI-GS: Interest-based Local Quality 3D Gaussian Splatting
por: Bui, Quoc-Anh, et al.
Publicado: (2025)
por: Bui, Quoc-Anh, et al.
Publicado: (2025)
ROI-NeRFs: Hi-Fi Visualization of Objects of Interest within a Scene by NeRFs Composition
por: Bui, Quoc-Anh, et al.
Publicado: (2025)
por: Bui, Quoc-Anh, et al.
Publicado: (2025)
Attention is also needed for form design
por: Sankar, B., et al.
Publicado: (2025)
por: Sankar, B., et al.
Publicado: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
por: Siddiqui, Yousuf Ahmed, et al.
Publicado: (2025)
por: Siddiqui, Yousuf Ahmed, et al.
Publicado: (2025)
Cost-Effective Attention Mechanisms for Low Resource Settings: Necessity & Sufficiency of Linear Transformations
por: Hosseini, Peyman, et al.
Publicado: (2024)
por: Hosseini, Peyman, et al.
Publicado: (2024)
Handling Out-of-Distribution Data: A Survey
por: Tamang, Lakpa, et al.
Publicado: (2025)
por: Tamang, Lakpa, et al.
Publicado: (2025)
Deep Spectral Meshes: Multi-Frequency Facial Mesh Processing with Graph Neural Networks
por: Kosk, Robert, et al.
Publicado: (2024)
por: Kosk, Robert, et al.
Publicado: (2024)
Group Inertial Poser: Multi-Person Pose and Global Translation from Sparse Inertial Sensors and Ultra-Wideband Ranging
por: Xue, Ying, et al.
Publicado: (2025)
por: Xue, Ying, et al.
Publicado: (2025)
Deep Learning Approaches for Medical Imaging Under Varying Degrees of Label Availability: A Comprehensive Survey
por: Ma, Siteng, et al.
Publicado: (2025)
por: Ma, Siteng, et al.
Publicado: (2025)
The Impact of Image Resolution on Face Detection: A Comparative Analysis of MTCNN, YOLOv XI and YOLOv XII models
por: Ömercikoğlu, Ahmet Can, et al.
Publicado: (2025)
por: Ömercikoğlu, Ahmet Can, et al.
Publicado: (2025)
HieraEdgeNet: A Multi-Scale Edge-Enhanced Framework for Automated Pollen Recognition
por: Long, Yuchong, et al.
Publicado: (2025)
por: Long, Yuchong, et al.
Publicado: (2025)
Ejemplares similares
-
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
por: Viveiros, André G., et al.
Publicado: (2025) -
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025) -
Depth Priors in Removal Neural Radiance Fields
por: Guo, Zhihao, et al.
Publicado: (2024) -
A Cost-Effective Eye-Tracker for Early Detection of Mild Cognitive Impairment
por: Greco, Danilo, et al.
Publicado: (2024) -
Few-Class Arena: A Benchmark for Efficient Selection of Vision Models and Dataset Difficulty Measurement
por: Cao, Bryan Bo, et al.
Publicado: (2024)