Lost in Translation: Modern Neural Networks Still Struggle With Small Realistic Image Transformations
Fuente:
arXiv
Guardado en:
| Autores principales: | Shifman, Ofir, Weiss, Yair |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Recurrent Neural Networks for Still Images
por: Dmitri, et al.
Publicado: (2024)
por: Dmitri, et al.
Publicado: (2024)
Intriguing Properties of Modern GANs
por: Friedman, Roy, et al.
Publicado: (2024)
por: Friedman, Roy, et al.
Publicado: (2024)
What do CNNs Learn in the First Layer and Why? A Linear Systems Perspective
por: Chowers, Rhea, et al.
Publicado: (2022)
por: Chowers, Rhea, et al.
Publicado: (2022)
Demystify Transformers & Convolutions in Modern Image Deep Networks
por: Hu, Xiaowei, et al.
Publicado: (2022)
por: Hu, Xiaowei, et al.
Publicado: (2022)
Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
por: Jang, Youngjoon, et al.
Publicado: (2025)
por: Jang, Youngjoon, et al.
Publicado: (2025)
Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues
por: Jang, Youngjoon, et al.
Publicado: (2025)
por: Jang, Youngjoon, et al.
Publicado: (2025)
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
por: Kargaran, Amir Hossein, et al.
Publicado: (2026)
por: Kargaran, Amir Hossein, et al.
Publicado: (2026)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
por: Deng, Ken, et al.
Publicado: (2026)
por: Deng, Ken, et al.
Publicado: (2026)
Distilled Pooling Transformer Encoder for Efficient Realistic Image Dehazing
por: Tran, Le-Anh, et al.
Publicado: (2024)
por: Tran, Le-Anh, et al.
Publicado: (2024)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
por: Kamoi, Ryo, et al.
Publicado: (2024)
por: Kamoi, Ryo, et al.
Publicado: (2024)
Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
por: Muralidharan, Srikanth, et al.
Publicado: (2025)
por: Muralidharan, Srikanth, et al.
Publicado: (2025)
Dual-domain Adaptation Networks for Realistic Image Super-resolution
por: Fang, Chaowei, et al.
Publicado: (2025)
por: Fang, Chaowei, et al.
Publicado: (2025)
LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
por: Kong, Fei, et al.
Publicado: (2025)
por: Kong, Fei, et al.
Publicado: (2025)
Image-to-Image Translation with Diffusion Transformers and CLIP-Based Image Conditioning
por: Zhu, Qiang, et al.
Publicado: (2025)
por: Zhu, Qiang, et al.
Publicado: (2025)
Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
por: Mazzucco, Silvio, et al.
Publicado: (2025)
por: Mazzucco, Silvio, et al.
Publicado: (2025)
Are you Struggling? Dataset and Baselines for Struggle Determination in Assembly Videos
por: Feng, Shijia, et al.
Publicado: (2024)
por: Feng, Shijia, et al.
Publicado: (2024)
Translating Imaging to Genomics: Leveraging Transformers for Predictive Modeling
por: Farooq, Aiman, et al.
Publicado: (2024)
por: Farooq, Aiman, et al.
Publicado: (2024)
Enhancing Image Restoration Transformer via Adaptive Translation Equivariance
por: Hu, JiaKui, et al.
Publicado: (2025)
por: Hu, JiaKui, et al.
Publicado: (2025)
Animate Your Motion: Turning Still Images into Dynamic Videos
por: Li, Mingxiao, et al.
Publicado: (2024)
por: Li, Mingxiao, et al.
Publicado: (2024)
Human Action Recognition in Still Images Using ConViT
por: Hosseyni, Seyed Rohollah, et al.
Publicado: (2023)
por: Hosseyni, Seyed Rohollah, et al.
Publicado: (2023)
Pairwise Alignment & Compatibility for Arbitrarily Irregular Image Fragments
por: Shahar, Ofir Itzhak, et al.
Publicado: (2025)
por: Shahar, Ofir Itzhak, et al.
Publicado: (2025)
EvoStruggle: A Dataset Capturing the Evolution of Struggle across Activities and Skill Levels
por: Feng, Shijia, et al.
Publicado: (2025)
por: Feng, Shijia, et al.
Publicado: (2025)
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval
por: Most, Alexander, et al.
Publicado: (2025)
por: Most, Alexander, et al.
Publicado: (2025)
LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer
por: Fei, Song, et al.
Publicado: (2025)
por: Fei, Song, et al.
Publicado: (2025)
Image Translation with Kernel Prediction Networks for Semantic Segmentation
por: Mata, Cristina, et al.
Publicado: (2025)
por: Mata, Cristina, et al.
Publicado: (2025)
Translation-Equivariance of Normalization Layers and Aliasing in Convolutional Neural Networks
por: Scanvic, Jérémy, et al.
Publicado: (2025)
por: Scanvic, Jérémy, et al.
Publicado: (2025)
LostPaw: Finding Lost Pets using a Contrastive Learning-based Transformer with Visual Input
por: Voinea, Andrei, et al.
Publicado: (2023)
por: Voinea, Andrei, et al.
Publicado: (2023)
Harmformer: Harmonic Networks Meet Transformers for Continuous Roto-Translation Equivariance
por: Karella, Tomáš, et al.
Publicado: (2024)
por: Karella, Tomáš, et al.
Publicado: (2024)
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
por: Gafni, Tomer, et al.
Publicado: (2025)
por: Gafni, Tomer, et al.
Publicado: (2025)
Refracting Reality: Generating Images with Realistic Transparent Objects
por: Yin, Yue, et al.
Publicado: (2025)
por: Yin, Yue, et al.
Publicado: (2025)
GSwap: Realistic Head Swapping with Dynamic Neural Gaussian Field
por: Zhou, Jingtao, et al.
Publicado: (2026)
por: Zhou, Jingtao, et al.
Publicado: (2026)
NBAvatar: Neural Billboards Avatars with Realistic Hand-Face Interaction
por: Svitov, David, et al.
Publicado: (2026)
por: Svitov, David, et al.
Publicado: (2026)
From Keypoints to Realism: A Realistic and Accurate Virtual Try-on Network from 2D Images
por: Toozandehjani, Maliheh, et al.
Publicado: (2025)
por: Toozandehjani, Maliheh, et al.
Publicado: (2025)
Translating Images to Road Network: A Sequence-to-Sequence Perspective
por: Lu, Jiachen, et al.
Publicado: (2024)
por: Lu, Jiachen, et al.
Publicado: (2024)
A Diffusion Model Translator for Efficient Image-to-Image Translation
por: Xia, Mengfei, et al.
Publicado: (2025)
por: Xia, Mengfei, et al.
Publicado: (2025)
Erased, But Not Forgotten: Erased Rectified Flow Transformers Still Remain Unsafe Under Concept Attack
por: Jiang, Nanxiang, et al.
Publicado: (2025)
por: Jiang, Nanxiang, et al.
Publicado: (2025)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
por: Saxon, Michael, et al.
Publicado: (2024)
por: Saxon, Michael, et al.
Publicado: (2024)
Lost in Tracking Translation: A Comprehensive Analysis of Visual SLAM in Human-Centered XR and IoT Ecosystems
por: Chandio, Yasra, et al.
Publicado: (2024)
por: Chandio, Yasra, et al.
Publicado: (2024)
Diff-Mosaic: Augmenting Realistic Representations in Infrared Small Target Detection via Diffusion Prior
por: Shi, Yukai, et al.
Publicado: (2024)
por: Shi, Yukai, et al.
Publicado: (2024)
Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
por: Hurt, J. Alex, et al.
Publicado: (2025)
por: Hurt, J. Alex, et al.
Publicado: (2025)
Ejemplares similares
-
Recurrent Neural Networks for Still Images
por: Dmitri, et al.
Publicado: (2024) -
Intriguing Properties of Modern GANs
por: Friedman, Roy, et al.
Publicado: (2024) -
What do CNNs Learn in the First Layer and Why? A Linear Systems Perspective
por: Chowers, Rhea, et al.
Publicado: (2022) -
Demystify Transformers & Convolutions in Modern Image Deep Networks
por: Hu, Xiaowei, et al.
Publicado: (2022) -
Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
por: Jang, Youngjoon, et al.
Publicado: (2025)