IPIS Lab and K-water partner to develop AI that detects sinkholes before they occur

IPIS Lab will build a sinkhole-detection system for K-water that aims to cut analysis time by 90% and reach 85% detection accuracy by March 2027. The company will generate 20,000 synthetic GPR training datasets to overcome the scarcity of real-world cavity data.

Categorized in: AI News IT and Development
Published on: Sep 07, 2026
IPIS Lab and K-water partner to develop AI that detects sinkholes before they occur

IPIS Lab, an AI synthetic-data company, will build a sinkhole-detection system for Korea Water Resources Corp. (K-water) that aims to cut analysis time by 90% and reach 85% detection accuracy by March 2027. The public-private partnership tackles a stubborn data bottleneck: underground cavities and sinkholes are rare events, so the real-world data needed to train detection models is critically scarce.

The company will not try to collect more field data. Instead, it will generate 20,000 synthetic GPR training datasets using a physics-based engine that simulates how ground-penetrating radar signals behave across different cavity sizes, depths, and soil conditions. K-water supplies the operational demand, real GPR survey records, and field-testing environments. IPIS Lab supplies the generation, labeling, and detection engines under a system it calls GroundSafe AI.

Why synthetic data changes the economics

GPR surveys have long depended on manual expert review. The process is slow, and inconsistent standards across institutions make it hard to scale inspections or decide where to excavate. For AI for IT & Development teams, the core constraint is familiar: low-frequency anomaly detection fails when training corpora are too thin. IPIS Lab sidesteps this by synthesizing signals from physical first principles rather than hunting for more real-world incidents.

The GroundSafe AI pipeline has three components. A generation engine creates GPR signal images from cavity and soil parameters. An automatic labeling engine tags abnormal signals and normalizes them across formats. A detection engine then classifies anomalies and outputs drilling-priority reports. The company said the two-dimensional image-detection problem mirrors work it has already commercialized - automatically identifying 23 defect types in sewer CCTV footage and producing synthetic data for defense drones.

Measurable targets, not lab experiments

The project runs for seven months through March 2027 under K-water's "2026 Public-Private Open Innovation" program. IPIS Lab has set concrete benchmarks: more than 20,000 training datasets, mean average precision (mAP) of at least 85%, and a 90% reduction in analysis time. The goal is to raise the technology readiness level to the stage immediately before commercial deployment. K-water will validate results inside its operations during the project, not after it ends.

IPIS Lab holds three related domestic patents and an official test report from the Korea Conformity Laboratories (KCL). The company was founded in September 2023 from technology developed at Chung-Ang University's Visual and Intelligent Systems Laboratory and has also been selected for a Ministry of SMEs and Startups R&D program focused on scenario-based synthetic data for defense and industrial AI.

From data scarcity to generation

"AI projects stop not because of the algorithm but when you realize there is no data to train it with," said Baek Jun-ki, CEO of IPIS Lab. "Our role is to turn the problem of collecting data into the problem of generating it." He added that the company intends to build an integrated safety solution that diagnoses both pipeline interiors and surrounding ground, with eventual expansion into underground safety assessments and digital water management for local governments.

Why this matters for IT and development professionals

The project is a worked example of synthetic data solving a real operational bottleneck, not a demo. For teams working on anomaly detection, predictive maintenance, or any domain where failure events are rare, the approach is directly transferable. The architecture - a physics-informed generator paired with automated labeling and detection - shows how to decouple model development from data collection when the thing you need to detect simply does not happen often enough. The Data Analysis pipeline here is worth studying: synthetic generation, standardized labeling, and a detection model that outputs ranked action items rather than raw classifications.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)