Hitachi unveils AI image tech that halves ViT processing time
KEY POINTS
- Hitachi develops selective region encoding technology to speed Vision Transformer processing of large image datasets
- Text-search evaluation halves average feature-generation time while maintaining search accuracy
- Technology targets manufacturing and maintenance image workflows and supports planned use in Lumada 3.0-related AI platforms

Hitachi Ltd. has developed a selective region encoding technology that speeds up AI processing of large volumes of image data by limiting Vision Transformer, or ViT, computation to selected parts of an image, the company said on August 25.
The technology is designed for image data generated in settings such as manufacturing and maintenance. ViT is a core architecture used in image AI models to generate feature data for tasks including image classification and search, but processing entire images requires handling large numbers of tokens tied to finely divided image regions, increasing processing time and demand for computing resources such as GPUs.
Hitachi analyzed which image regions contribute to maintaining accuracy and found that tokens corresponding to regions containing the target object were important. It also found that, across many images, the AI model concentrated attention on specific fixed positions, and that tokens corresponding to those fixed regions also helped preserve accuracy.
Based on those findings, Hitachi adopted a method that processes only tokens tied to detected object regions and fixed regions. In an evaluation using generated features for a search task targeting text within images, the company confirmed that the technology halved the average processing time required for feature generation while maintaining search accuracy, compared with a conventional method that processes the entire image. The test used the TextOCR dataset under evaluation conditions designed for searches related to text appearing in images, and the effect may vary depending on the type of object, how it appears in an image and the GPU used.
Hitachi said the technology could help accelerate AI processing for large image datasets while reducing computing resource requirements, including GPUs. Part of the research has been published in The IEICE Transactions on Information and Systems.
The company plans to verify the technology's value in practical use cases through collaborations with customers and consider applying it to AI platforms that support the use of large volumes of image and video data. Hitachi also positioned the work as one of the technologies supporting stronger Lumada 3.0 offerings, its digital services platform.
Originally published on jp.ibtimes.com
© Copyright 2026 IBTimes JP. All rights reserved.



















