DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model
Abstract
Significant progress has been made in the field of Instruction-based Image Editing Models (IIEMs). However, while these models demonstrate plausible adherence to instructions and strong reasoning ability on current benchmarks, their ability to edit small objects remains underexplored, despite its importance for precise local editing and refining details in both real and generated images. In this paper, we introduce DeepLookEditBench (DLEBench), the first benchmark dedicated to assessing the abilities of IIEMs in editing small-scale objects. Specifically, we construct a challenging testbed comprising 1889 samples across seven instruction types. In these samples, target objects occupy only 1%-10% of the image area, covering complex scenarios such as partial occlusion and multi-object editing. To ensure robust evaluation on this benchmark, we propose an evaluation protocol with refined score rubrics to minimize subjectivity and ambiguity in two criteria: Instruction Following and Visual Consistency. This protocol also introduces a dual-mode evaluation framework (Tool-driven and Oracle-guided Modes) addressing the misalignment between LMM-as-a-Judge and human judgments on DLEBench. Empirical results on 10 IIEMs reveal significant performance gaps in small-scale object editing, highlighting the need for specialized benchmarks to advance this ability.
Lay Summary
Instruction-based image editing systems can now handle many user requests, but we still know surprisingly little about how well they handle very small objects in images. This matters because many practical edits, such as correcting details or refining local regions, depend on changing tiny objects accurately without harming the rest of the picture. To study this problem, we introduce DLEBench, the first benchmark designed specifically to test small-object editing in these systems. Our benchmark contains 1,889 examples covering seven types of editing instructions, including difficult cases in which small objects are partially occluded or appear alongside multiple other objects. We also design a more rigorous evaluation method to assess whether a model both follows the instructions and preserves the image's visual quality. Because existing vision-language judges often miss subtle changes in tiny regions, our method includes two evaluation modes to better align automatic scoring with human judgment. Tests on 10 image editing models show clear weaknesses on this task. Our results suggest that progress on general benchmarks does not yet translate to reliable small-object editing.