Modeling Image Tone Dichotomy with the Power Function
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Martinez, Axel, Olague, Gustavo, Hernandez, Emilio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Analytical-Heuristic Modeling and Optimization for Low-Light Image Enhancement
von: Martinez, Axel, et al.
Veröffentlicht: (2024)
von: Martinez, Axel, et al.
Veröffentlicht: (2024)
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
von: Yang, Sheng, et al.
Veröffentlicht: (2025)
von: Yang, Sheng, et al.
Veröffentlicht: (2025)
Complete Gaussian Splats from a Single Image with Denoising Diffusion Models
von: Liao, Ziwei, et al.
Veröffentlicht: (2025)
von: Liao, Ziwei, et al.
Veröffentlicht: (2025)
PALMS+: Modular Image-Based Floor Plan Localization Leveraging Depth Foundation Model
von: Cheng, Yunqian, et al.
Veröffentlicht: (2025)
von: Cheng, Yunqian, et al.
Veröffentlicht: (2025)
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
Efficient Construction of Implicit Surface Models From a Single Image for Motion Generation
von: Chu, Wei-Teng, et al.
Veröffentlicht: (2025)
von: Chu, Wei-Teng, et al.
Veröffentlicht: (2025)
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
von: Singh, Ishika, et al.
Veröffentlicht: (2025)
von: Singh, Ishika, et al.
Veröffentlicht: (2025)
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization
von: Kawaharazuka, Kento, et al.
Veröffentlicht: (2024)
von: Kawaharazuka, Kento, et al.
Veröffentlicht: (2024)
DARTS: A Drone-Based AI-Powered Real-Time Traffic Incident Detection System
von: Li, Bai, et al.
Veröffentlicht: (2025)
von: Li, Bai, et al.
Veröffentlicht: (2025)
DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features
von: Wang, Letian, et al.
Veröffentlicht: (2024)
von: Wang, Letian, et al.
Veröffentlicht: (2024)
Tether: Autonomous Functional Play with Correspondence-Driven Trajectory Warping
von: Liang, William, et al.
Veröffentlicht: (2026)
von: Liang, William, et al.
Veröffentlicht: (2026)
A Deep Learning-based Pest Insect Monitoring System for Ultra-low Power Pocket-sized Drones
von: Crupi, Luca, et al.
Veröffentlicht: (2024)
von: Crupi, Luca, et al.
Veröffentlicht: (2024)
SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization
von: Chen, Posheng, et al.
Veröffentlicht: (2026)
von: Chen, Posheng, et al.
Veröffentlicht: (2026)
FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction
von: Rotondi, Dennis, et al.
Veröffentlicht: (2025)
von: Rotondi, Dennis, et al.
Veröffentlicht: (2025)
Multimodal Object Detection using Depth and Image Data for Manufacturing Parts
von: Mahjourian, Nazanin, et al.
Veröffentlicht: (2024)
von: Mahjourian, Nazanin, et al.
Veröffentlicht: (2024)
Optimization of Autonomous Driving Image Detection Based on RFAConv and Triplet Attention
von: Ling, Zhipeng, et al.
Veröffentlicht: (2024)
von: Ling, Zhipeng, et al.
Veröffentlicht: (2024)
RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment
von: Zhuang, Qiyuan, et al.
Veröffentlicht: (2026)
von: Zhuang, Qiyuan, et al.
Veröffentlicht: (2026)
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
von: Wang, Yihao, et al.
Veröffentlicht: (2025)
von: Wang, Yihao, et al.
Veröffentlicht: (2025)
FUNCanon: Learning Pose-Aware Action Primitives via Functional Object Canonicalization for Generalizable Robotic Manipulation
von: Xu, Hongli, et al.
Veröffentlicht: (2025)
von: Xu, Hongli, et al.
Veröffentlicht: (2025)
MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence
von: Tang, Chao, et al.
Veröffentlicht: (2025)
von: Tang, Chao, et al.
Veröffentlicht: (2025)
Multi-Object Tracking based on Imaging Radar 3D Object Detection
von: Palmer, Patrick, et al.
Veröffentlicht: (2024)
von: Palmer, Patrick, et al.
Veröffentlicht: (2024)
ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition
von: Xie, Weidong, et al.
Veröffentlicht: (2024)
von: Xie, Weidong, et al.
Veröffentlicht: (2024)
SynDroneVision: A Synthetic Dataset for Image-Based Drone Detection
von: Lenhard, Tamara R., et al.
Veröffentlicht: (2024)
von: Lenhard, Tamara R., et al.
Veröffentlicht: (2024)
Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement
von: Feng, Zhicheng, et al.
Veröffentlicht: (2025)
von: Feng, Zhicheng, et al.
Veröffentlicht: (2025)
Gradient-Guided Parameter Mask for Multi-Scenario Image Restoration Under Adverse Weather
von: Guo, Jilong, et al.
Veröffentlicht: (2024)
von: Guo, Jilong, et al.
Veröffentlicht: (2024)
CoFiI2P: Coarse-to-Fine Correspondences for Image-to-Point Cloud Registration
von: Kang, Shuhao, et al.
Veröffentlicht: (2023)
von: Kang, Shuhao, et al.
Veröffentlicht: (2023)
ParkingE2E: Camera-based End-to-end Parking Network, from Images to Planning
von: Li, Changze, et al.
Veröffentlicht: (2024)
von: Li, Changze, et al.
Veröffentlicht: (2024)
Sim2real Image Translation Enables Viewpoint-Robust Policies from Fixed-Camera Datasets
von: Coholich, Jeremiah, et al.
Veröffentlicht: (2026)
von: Coholich, Jeremiah, et al.
Veröffentlicht: (2026)
MetricGold: Leveraging Text-To-Image Latent Diffusion Models for Metric Depth Estimation
von: Shah, Ansh, et al.
Veröffentlicht: (2024)
von: Shah, Ansh, et al.
Veröffentlicht: (2024)
Cycle-Correspondence Loss: Learning Dense View-Invariant Visual Features from Unlabeled and Unordered RGB Images
von: Adrian, David B., et al.
Veröffentlicht: (2024)
von: Adrian, David B., et al.
Veröffentlicht: (2024)
Depth-aware Fusion Method based on Image and 4D Radar Spectrum for 3D Object Detection
von: Sun, Yue, et al.
Veröffentlicht: (2025)
von: Sun, Yue, et al.
Veröffentlicht: (2025)
MTR-VP: Towards End-to-End Trajectory Planning through Context-Driven Image Encoding and Multiple Trajectory Prediction
von: Keskar, Maitrayee, et al.
Veröffentlicht: (2025)
von: Keskar, Maitrayee, et al.
Veröffentlicht: (2025)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control
von: Chen, Yuxin, et al.
Veröffentlicht: (2025)
von: Chen, Yuxin, et al.
Veröffentlicht: (2025)
SEAL: Vision-Language Model-Based Safe End-to-End Cooperative Autonomous Driving with Adaptive Long-Tail Modeling
von: You, Junwei, et al.
Veröffentlicht: (2025)
von: You, Junwei, et al.
Veröffentlicht: (2025)
Visual SLAMMOT Considering Multiple Motion Models
von: Tian, Peilin, et al.
Veröffentlicht: (2024)
von: Tian, Peilin, et al.
Veröffentlicht: (2024)
Universal Actions for Enhanced Embodied Foundation Models
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
Language-Conditioned World Modeling for Visual Navigation
von: Dong, Yifei, et al.
Veröffentlicht: (2026)
von: Dong, Yifei, et al.
Veröffentlicht: (2026)
Rethinking Video Generation Model for the Embodied World
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
MemoNav: Working Memory Model for Visual Navigation
von: Li, Hongxin, et al.
Veröffentlicht: (2024)
von: Li, Hongxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Analytical-Heuristic Modeling and Optimization for Low-Light Image Enhancement
von: Martinez, Axel, et al.
Veröffentlicht: (2024) -
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
von: Yang, Sheng, et al.
Veröffentlicht: (2025) -
Complete Gaussian Splats from a Single Image with Denoising Diffusion Models
von: Liao, Ziwei, et al.
Veröffentlicht: (2025) -
PALMS+: Modular Image-Based Floor Plan Localization Leveraging Depth Foundation Model
von: Cheng, Yunqian, et al.
Veröffentlicht: (2025) -
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)