ForensiText: ROI-Guided Multimodal Forensics for Scene Text Tampering Detection and Localization
Abhineet Kumar Pandey; Ming-Ching Chang
22nd International Conference On Advanced Visual And Signal-Based Systems (AVSS)Scene text manipulation is a subtle yet consequential form of visual tampering, where small edits to signs, receipts, labels, or prices can significantly alter image meaning while leaving minimal forensic traces. Existing forensic methods either analyze the full image, diluting evidence from small text regions, or adopt OCR-centric approaches that prioritize textual content over visual authenticity. This paper presents ForensiText, an ROI-guided forensic feature learning framework for scene text tampering detection and localization. Given an RGB image, a scene-text detector estimates text-bearing regions containing both pristine and manipulated text; ground-truth manipulation masks are used only for pixel-level supervision and evaluation. Within the detected text ROIs, the framework extracts complementary forensic cues from
ROI-gated RGB appearance, SRM residuals, CFA inconsistencies, JPEG error-level analysis, and optional Noiseprint++ features. These streams are fused with an explicit text-ROI channel and processed by a compact dual-head encoder-decoder to jointly predict an image-level tampering score and a pixel-level manipulation mask. By concentrating forensic reasoning on automatically detected text regions, ForensiText preserves fine-grained localization while suppressing irrelevant background noise. Under a detector-guided text-ROI evaluation protocol, the method achieves strong pixel-level localization performance across multiple benchmarks, with Dice scores of 0.8684 on RealTextManipulation, 0.8364 on OSTF, 0.9843 on Tampered-IC13, and 0.8269 on TextSleuth. These results demonstrate the value of combining text-region priors with low-level forensic evidence for scene text manipulation detection.