
ULSAN — A Korean research team has identified structural flaws in the global standard used to assess whether artificial intelligence identifies objects based on their outline, or shape, or on their surface patterns, or texture. The team's new evaluation framework, which overhauls a scoring method that distorted AI performance by relying on simple ratio calculations, is drawing attention in the field.
The Ulsan National Institute of Science and Technology (UNIST) said on the 1st that a team led by Professor Yoo Jae-jun of its Graduate School of Artificial Intelligence had demonstrated the flaws in "cue-conflict," the existing method for evaluating AI vision, and built a replacement benchmark called REFINED-BIAS.
Researchers have measured how AI recognizes objects using cue-conflict, devised in 2019 by a team at the University of Tübingen in Germany. The method presents a composite image — a cat's body overlaid with elephant skin texture, for example — and checks which cue the model answers with. Early results suggested that AI, unlike humans, leans toward texture, and the prevailing view held that models should be trained to prioritize shape as humans do. Later studies produced conflicting findings on the correlation between shape preference and model performance, keeping the debate alive.
The team traced the confusion to limitations within the benchmark itself. The existing method measured only the ratio of an AI's choices, failing to capture how well it actually uses each cue. A high-performing model that correctly identified shape 80 times out of 100 questions and a low-performing model that managed only eight both received a shape preference score of 80 percent, as long as their texture recognition rates were low. The method could not distinguish a weak model whose figure rose only because it failed to recognize texture at all. The team also found scoring flaws in which shape and texture were insufficiently separated within composite images, and in which a model's top answer was discarded for falling outside a preset list of candidates, letting its second choice pass as correct.
REFINED-BIAS shifts the scoring system from ratios to individual scores measuring how sensitively a model uses each of the two cues. The team selected 10 categories with distinctive shapes, such as an hourglass, and 10 with distinctive textures, such as a zebra, and built 6,000 new images. For shape evaluation, surface patterns were erased to leave only outlines; for texture evaluation, images were cut and rearranged to hide form. The evaluation range was also widened from a set list of candidates to every option a model can produce.
Testing under the new standard showed that stronger models do not lean toward one cue but actively use both shape and texture. What separated models was not the direction of their preference but the total amount of visual information they drew on. The results also clearly confirmed that recent models built on transformer architectures or trained jointly on text and images are indeed better at recognizing shape.
"A benchmark is more than a way to measure AI performance — it is a reference point that sets the direction of future development," Yoo said. "This research will play a central role in improving model architectures and training methods by precisely diagnosing how visual AI uses information."
The paper, co-first-authored by researchers Kim Beom-jun and Lee Seung-a, was selected as a Spotlight paper at the European Conference on Computer Vision (ECCV 2026), the leading conference in computer vision, placing it in the top 1.3 percent of 10,473 submissions.






