Research output

Publications and patents

Records migrated from the prior portfolio. Incomplete bibliographic fields remain visibly marked for review rather than being inferred.

Peer-reviewed publications

Conference paper2026

ForensiText: ROI-Guided Multimodal Forensics for Scene Text Tampering Detection and Localization

Abhineet Kumar Pandey; Ming-Ching Chang

22nd International Conference On Advanced Visual And Signal-Based Systems (AVSS)

Scene text manipulation is a subtle yet consequential form of visual tampering, where small edits to signs, receipts, labels, or prices can significantly alter image meaning while leaving minimal forensic traces. Existing forensic methods either analyze the full image, diluting evidence from small text regions, or adopt OCR-centric approaches that prioritize textual content over visual authenticity. This paper presents ForensiText, an ROI-guided forensic feature learning framework for scene text tampering detection and localization. Given an RGB image, a scene-text detector estimates text-bearing regions containing both pristine and manipulated text; ground-truth manipulation masks are used only for pixel-level supervision and evaluation. Within the detected text ROIs, the framework extracts complementary forensic cues from

ROI-gated RGB appearance, SRM residuals, CFA inconsistencies, JPEG error-level analysis, and optional Noiseprint++ features. These streams are fused with an explicit text-ROI channel and processed by a compact dual-head encoder-decoder to jointly predict an image-level tampering score and a pixel-level manipulation mask. By concentrating forensic reasoning on automatically detected text regions, ForensiText preserves fine-grained localization while suppressing irrelevant background noise. Under a detector-guided text-ROI evaluation protocol, the method achieves strong pixel-level localization performance across multiple benchmarks, with Dice scores of 0.8684 on RealTextManipulation, 0.8364 on OSTF, 0.9843 on Tampered-IC13, and 0.8269 on TextSleuth. These results demonstrate the value of combining text-region priors with low-level forensic evidence for scene text manipulation detection.

Conference paper2025

Integrating manual preprocessing with automated feature extraction for improved rodent seizure classification

An Yu, Mannut Singh, Abhineet Pandey, Elizabeth Dybas, Aditya Agarwal, Yifan Kao, Guangliang Zhao, Tzu-Jen Kao, Xin Li, Damian S Shin, Ming-Ching Chang

Epilepsy and Behavior
Artificial IntelligenceSeizure Stage RecognitionRodent ModelEpilepsyBehavior AnalysisVideo AnalyticsRacine ScaleSkeleton Keypoints

Hypothesis/Objective

Approach/Method

Results

Conclusions

Rodent models of epilepsy can help with the search for more effective drug candidates or neuromodulatory therapies. Yet, preclinical screening of candidate options for anti-epileptic drugs (AED) using rodent models may require hours of video monitoring. Data processing is also time-consuming, subjective, and error-prone. This study aims to develop an AI-enabled quantitative analysis of rodent behavior, including epilepsy stage classification.We leveraged deep learning and computer vision techniques to develop a semi-automatic pipeline and framework for animal seizure detection and recognition, which requires manual preprocessing of the dataset. Our hybrid approach combines model-based and data-driven methods but is dependent on manually preprocessed and segmented video clips to facilitate the automatic classification of epilepsy stages.We collected two datasets comprising rat skeleton keypoints and seizure behavior videos in the lab. The proposed method, PoseC3D, for rat seizure stage classification of the collected database achieved an accuracy between 64.7–90.3 % when tested on four different seizure phenotypes using the Racine classification scale.This study demonstrates the feasibility of video-based seizure stage detection and classification for rodent models of temporal lobe seizures using a semi-automatic pipeline that requires manual preprocessing of data. However, our method is not capable of fully automated seizure detection and has not been tested on unseen animals, which limits its generalizability and applicability for broader use. Despite these limitations, the approach underscores our ability to undertake quantitative analysis of rodent behavior, which can also support other studies of animal behavior involving motor functions and future considerations for non-motor symptomology such as mood disorders.

Conference paper2024

TextSleuth: A New Dataset and Baseline for Scene Text Manipulation Detection

Abhineet Kumar Pandey; Ming-Ching Chang; Xin Li

2024 IEEE 7th International Conference on Multimedia Information Processing and Retrieval (MIPR)
Scene TextMedia Forensics

With the rise of digital content on social media and the advancement of image editing tools, tampering with scene text has become a serious concern. Scene text manipulation detection (STMD) is a kind of image manipulation detection (IMD) with focus on the tampering of scene text pixels, which is crucial for image content integrity and media forensics. In this paper, we present TextSleuth, a novel benchmark dataset specifically designed for STMD, by integrating three public datasets with newly introduced manipulation and annotations. We introduce professional edits on the Total-Text dataset (~1K images) with four levels of manipulated region perceptibility, and a large synthetic manipulation set (858K images) on the SynthText dataset, as well the integration of the Tampered-IC13 dataset (378 images). We established a new STMD baseline based on TextSleuth using MMFusion-IML, the state-of-the-art image manipulation detection model. We performed extensive experiments, reporting the AUC from ROC analysis and the balanced accuracy (bACC) metrics to maintain a balanced performance evaluation. The MMFusion-IML baseline achieves 0.641 AUC and 0.588 bACC on the Total-Text subset. In comparison, it achieves 0.89 AUC and 0.8272 bACC on the Tampered-IC13 subset. This showcases the real-world STMD challenges reflected in our new dataset. TextSleuth is a valuable resource for future research in scene text manipulation detection and forensics. The dataset is available at https://github.com/abhineet-pandey/Text-Sleuth.

Conference paper2021

FlagDetSeg: Multi-Nation Flag Detection and Segmentation in the Wild

Shou-Fang Wu; Ming-Ching Chang; Siwei Lyu; Cheng-Shih Wong; Abhineet Kumar Pandey; Po-Chi Su

2021 17th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS)
DetectionSegmentation

We present a simple and effective flag detection approach for multi-nation flag instance segmentation in-the-wild based on data augmentation and Mask-RCNN PointRend. To the best of our knowledge, this is the first multi-nation flag detection work incorporating recent deep object detection with code and dataset that will be released for public use. Flag images with binary segmentation are collected from public domain including the Open Image V6 and annotated for up to 225 countries. Additional flag images are generated from template flag images with cropping, warping, masking, and color adaption to hallucinate realistic-looking flag images for training and testing. Data augmentation is performed by fusing and transforming the segmented flags on top of natural image backgrounds to synthesize new images. To cope with the large variability of flags with the lack of authentic annotated flags, we combine the trained binary Mask-RCNN segmentation weights with the new multi-nation classifier for fine-tuning. For evaluation, the proposed model is compared with other popular detectors and instance segmentation methods including YOLACT++. Results show the efficacy of the proposed approach

Conference paper2021

A Video Analytic System for Rail Crossing Point Protection

Guangliang Zhao; Abhineet Kumar Pandey; Ming-Ching Chang; Siwei Lyu

2021 17th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS)
Video AnalyticsTrackingRail Safety

With the rise of AI deep learning, video surveillance based on deep neural networks can provide real-time detection and tracking of vehicles and pedestrians. We present a video analytic system for monitoring railway crossing and providing security protection for rail intersections. Our system can automatically determine the rail-crossing gate status via visual detection and analyze traffic by detecting and tracking passing vehicles, thus to oversee a set of rail-transportation related safety events. Assuming a fixed camera view, each gate RoI can be manually annotated once for each site during system setup, and then gate status can be automatically detected afterwards. Vehicles are detected using YOLOv4 and multi-target tracking is performed using DeepSORT. Safety-related events including trespassing are continuously monitored using rule-based triggering. Experimental evaluation is performed on a Youtube rail crossing dataset as well as a private dataset. On the private dataset of 76 total minutes from 38 videos, our system can successfully detect all 56 events out of 58 annotated events. On the public dataset of 14.21 hrs of videos, it detects 58 out of 62 events.

Conference paper2020

Railcar Detection, Identification and Tracking for Rail Yard Management

Ming-Ching Chang; Guangliang Zhao; Abhineet Kumar Pandey; Andrew Pulver; Peter Tu

2020 IEEE International Conference on Image Processing (ICIP)
DetectionTrackingRe-identification

We present a video analytics system combining railcar detection, classification, Federal Railroad Admin. (FRA) text identification, and logo detection into a system for locomotive transportation and yard management. Existing RFID-based systems are limited by sensor deployment and cannot visually identify railcars when they are away. As there are typically tens of tracks and hundreds of railcars in a yard, an automatic vision system is desirable. The proposed AI system is developed for autonomous yard inventory checking, such that the arrival, departure, and movement of individual railcars can be automatically monitored and managed in the facility. Our system consists of multiple cameras with edge computing devices installed at check points (track entrances and branches), such that visual detection and tracking of railcars can be performed and meta-data can be exchanged. After knowing the railcar locations and types, scene text detection is performed to search and recognize FRA ID markings and logos that can uniquely identify each railcar. Information fusion a database in the central hub can further improve railcar identification and reduce errors. Early results on real-world field collected data demonstrate the efficacy of the proposed approach.

Patents

WO2025165595A12025

Depth-assisted sample container characterization

Published

Patent record →

An automated diagnostic analysis system includes an input module that receives at least one sample container holder that includes one or more sample containers to be processed by the system. The input module includes an imaging sensor for capturing at a tilted angle one or more images of the one or more sample containers. The system also includes a computer processor executing a trained machine learning model for analyzing the one or more images to obtain or refine three-dimensional depth information from the images to determine physical characteristics of each sample container. Methods of operating an automated diagnostic analysis system are also provided, as are other aspects.

US20250271454A12023

Sample handlers of diagnostic laboratory analyzers and methods of use

Publication

Patent record →

A sample handler of a diagnostic laboratory system includes a plurality of holding locations configured to receive sample containers. An imaging device is movable within the sample handler and is configured to capture images of the holding locations and sample containers received therein. A controller is configured to generate instructions that cause the imaging device to move within the sample handler and capture images. A classification algorithm is implemented in computer code, and includes a trained model configured to classify objects in the captured images. Other sample handlers and methods of handling sample containers are disclosed.