Complete AI Training

AI news ·

Tencent's AIG joins ClawScan to review every skill and plugin on ClawHub

Tencent's AIG scanner now runs alongside NVIDIA's SkillSpector on every ClawHub upload, catching 98.6% of malicious cases across a 556-sample benchmark.

Share

Tencent's AI-Infra-Guard (AIG) has been integrated into ClawScan, the open-source command-line tool that handles security review on ClawHub. Every skill and plugin uploaded to the platform now runs through AIG as part of its security review, adding a second independent scanner alongside NVIDIA's SkillSpector to catch a broader range of risks.

How the dual-scanner review works

ClawScan runs AIG and SkillSpector independently on each submission. An AI judge then reviews both scanners' findings alongside the uploaded files and makes a final security assessment. The approach preserves the different perspectives each scanner brings - a design choice informed by ClawHub's Security Signals paper, which showed how differently scanners can interpret the same skill.

Contributors can test different scanners, AI models, and review instructions against the same benchmark to measure whether a change improves results. This keeps the review pipeline open to experimentation and external validation.

What AIG brings to the pipeline

Tencent describes AIG's skill scanner as "an LLM-driven multi-stage code audit and vulnerability review pipeline." It traces the relationship between a skill's instructions and the scripts, dependencies, and data flows behind them. AIG reviews nine categories of risk: instruction hijacking, memory poisoning, remote payload execution, unauthorized access, persistence, insecure dependencies, and three additional vulnerability classes.

The scanner's vulnerability-pattern checks surface in ClawHub's security audit view, giving reviewers direct visibility into flagged risks rather than a pass/fail verdict.

Benchmarking against SkillTrustBench

Tencent and ClawHub evaluated the combined system on a fixed 556-case subset of SkillTrustBench, a public benchmark developed by Tencent and the Chinese University of Hong Kong, Shenzhen. The benchmark includes benign, suspicious, and malicious skills across all nine risk categories. It tests whether scanners catch risky behavior, avoid flagging legitimate skills, and distinguish security flaws from malicious intent.

Across those 556 cases, ClawScan matched 86.9% of the benchmark's labels and correctly classified 98.6% of malicious cases. Tencent found that AIG and SkillSpector surfaced different risks even when using the same AI model, confirming the value of running both scanners in parallel.

A shared feedback loop

The two teams now exchange anonymized cases where the scanners disagree, along with confirmed false positives. These examples feed into regression tests and improvements to AIG's detection rules and review process. Changes get evaluated through ClawScan, creating a feedback loop that sharpens both projects over time.

Keeping the work open means contributors can build on the integration and bring their own evidence to the next round of improvements.

Why this matters for IT and development teams

For teams building or consuming community-contributed skills and plugins, the dual-scanner approach reduces the chance that a single tool's blind spot lets a malicious submission through. The 98.6% classification rate on malicious cases is a strong signal, but the real value is in the architecture: two independent scanners with different detection philosophies, reviewed by an AI judge, on an open-source pipeline that anyone can benchmark. Security engineers and DevSecOps leads evaluating supply-chain risks in plugin ecosystems can study the ClawScan model as a reference for their own review workflows - and professionals looking to build expertise in this area can explore structured AI Safety Engineering Courses and AI Security Analytics Courses that cover vulnerability detection and automated review pipelines.

Share