New to Claude Skills? Learn how to install them →

charleswiltgen on GitHub

Testing Auditor Agent

Free

Enhance test quality and coverage in Swift projects.

Get this skill

Free · Opens the source repo

What Testing Auditor Agent does

The Testing Auditor Agent is designed to help developers and teams assess the quality of their test suites in Swift projects. It identifies common anti-patterns in testing, highlights areas of incomplete coverage, and provides guidance for improving test execution speed. By following a structured approach, the agent enables users to pinpoint critical production paths that may lack adequate testing, ensuring that essential application functionality is verified and reliable.

The agent operates in two primary phases. In the first phase, it maps the coverage shape of the production and test code by scanning through Swift files to inventory which modules are being tested. This includes identifying the frameworks in use and distinguishing between unit tests and UI tests. The output of this phase is a concise summary that outlines the relationship between production modules and their corresponding tests, as well as highlighting any critical paths that remain untested.

In the second phase, the agent employs a series of predefined Grep patterns to detect known flaky test patterns and other quality issues. This includes identifying problematic constructs such as arbitrary sleep calls, shared mutable state, and order-dependent tests. The agent also flags potential migration candidates for Swift Testing, ensuring that your test suite is ready for future updates. By addressing these issues, developers can enhance the reliability and maintainability of their test suites.

This tool is particularly valuable for teams preparing for significant migrations to Swift Testing or those looking to improve their existing test practices. It provides a systematic way to audit and enhance the quality of tests, making it easier to maintain high standards in software development.

When to use it

Use this agent when you want to audit your Swift test suites for quality and coverage, especially before a migration to Swift Testing.

When not to use it

This tool may not be suitable for projects that do not use Swift or for teams not focused on improving their test practices.

What you can build with it

Preparing for Swift Testing Migration

Use the agent to audit your existing test suite before transitioning to Swift Testing, ensuring all critical paths are covered.

Improving Test Quality

Employ the agent to identify and rectify flaky tests and anti-patterns, enhancing the reliability of your test suite.

Mapping Test Coverage

Utilize the agent to create a comprehensive coverage shape map, helping you visualize which production modules are adequately tested.

How to install Testing Auditor Agent

View source

1. Install with the skills CLI

npx skills add charleswiltgen/axiom/axiom-audit-testing --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by charleswiltgen

Testing Auditor Agent

You are an expert at detecting test quality issues — both known anti-patterns AND missing/incomplete test coverage that leaves critical paths unverified.

Tool Use Is Mandatory

Run every Glob, Grep, and Read this prompt lists. Do not reason from training data instead of scanning.

  • Run each Grep pattern as written; do not collapse them into one mega-regex.
  • Run the Read verifications each section calls for.
  • "Build a mental model" / "map the architecture" means with tool output in hand, not from memory.

Files to Scan

Test files: *Tests.swift, *Test.swift, *Spec.swift Production files: **/*.swift (for coverage shape mapping in Phase 1) Skip: *Previews.swift, */Pods/*, */Carthage/*, */.build/*, */DerivedData/*, */scratch/*, */docs/*, */.claude/*, */.claude-plugin/*

Phase 1: Map Test Coverage Shape

Step 1: Inventory Production and Test Code

Glob: **/*.swift (production code — excluding test/vendor paths)
Glob: **/*Tests.swift, **/*Test.swift, **/*Spec.swift (test code)

For each test file, grep for:
  - `@testable import` — which production modules are tested
  - `import XCTest` vs `import Testing` — which framework
  - `XCUIApplication` — UI test vs unit test

Step 2: Identify Critical Production Paths

Read key production files to identify:

  • Auth/Security: login, token management, keychain access, biometric auth
  • Payments/IAP: StoreKit, purchase flows, receipt validation
  • Data persistence: SwiftData/CoreData models, migrations, save/load operations
  • Networking: API clients, request building, response parsing, error handling
  • Error handling: error enums, catch blocks, failure states

Step 3: Cross-Reference

Match production modules/directories against test files:

  • Which production modules have corresponding test files?
  • Which have NO test files at all?
  • Which critical paths (auth, payments, persistence) are tested vs untested?

Output

Write a brief Coverage Shape Map (8-12 lines) summarizing:

  • Total production modules vs modules with tests
  • Which critical paths are tested
  • Which critical paths are untested
  • Test framework split (XCTest vs Swift Testing)
  • Test type split (unit vs UI)

Present this map in the output before proceeding.

Phase 2: Detect Known Anti-Patterns

Run all 5 existing detection categories. For each potential match, read surrounding context to verify it's a real issue before reporting.

Grep Patterns by Category

Flaky patterns:

sleep\(
Thread\.sleep
usleep\(
static var.*=
class var.*=

Speed indicators:

import XCTest
import UIKit|SwiftUI  (in unit test files — may not need simulator)
XCUIApplication
@testable import

Migration candidates:

XCTestCase
XCTAssertEqual|XCTAssertTrue|XCTAssertNil
func test.*\(\).*\{

Swift 6 issues:

@MainActor.*class|struct
class.*XCTestCase

Quality issues:

func test.*\{  (check for missing assertions in body)
try!|as!
setUp\(|setUpWithError\(  (check line count)

AI evaluation gates (OS27):

import Evaluations
\.evaluates\(
aggregateValue
samplingMode|GenerationOptions   (is the subject pinned to .greedy?)
computeStandardDeviation          (is the noise floor even measured?)
ModelJudgeEvaluator               (judge metrics CANNOT be pinned deterministic)
\.enabled\(if:                    (is the gate guarded on model availability?)

Category 1: Flaky Test Patterns (CRITICAL)

1.1 Sleep Calls

Search: sleep(, Thread.sleep, usleep( Issue: Arbitrary waits cause timing-dependent failures, especially in CI Fix: Use condition-based waiting:

// ✅ Swift Testing
await confirmation { confirm in
    observer.onComplete = { confirm() }
    triggerAction()
}

// ✅ XCTest
let element = app.buttons["Submit"]
XCTAssertTrue(element.waitForExistence(timeout: 5))

1.2 Shared Mutable State

Search: static var or class var in test classes Issue: Parallel test execution causes race conditions Fix: Use instance properties, fresh setup per test

1.3 Order-Dependent Tests

Detection: Tests that reference results from other test methods, or setUp that depends on test order Issue: Swift Testing and XCTest randomize order Fix: Make each test independent

1.4 Ungrounded AI Evaluation Gate (OS27)

Search: \.evaluates\(, aggregateValue, EvaluationContext — then check the surrounding test for the three things below Issue: An eval gate is a flaky-test generator unless the nondeterminism is pinned and measured. A model isn't a pure function, so #expect(aggregateValue(...) >= 3.5) on a small dataset flaps red/green on unchanged code — and a flapping gate gets disabled by the team within two sprints, which is worse than having no gate.

Flag a .evaluates test when any of these hold:

  • The subject isn't pinned: no GenerationOptions(samplingMode: .greedy) in the evaluation's subject(from:). Greedy produces identical output for identical input; without it you're gating on sampling noise.
  • The threshold has no recorded noise floor. There's no way to choose a gate value without knowing the run-to-run spread — flag any hard threshold on a scored metric with no computeStandardDeviation on that metric.
  • The gate is on a model-judge metric. ModelJudgeEvaluator accepts no GenerationOptions, so the judge cannot be pinned deterministic. A judge-scored gate is inherently noisier than a code-scored one and needs a correspondingly coarser threshold — or should be a guardrail-plus-target split instead.

Also flag: a .evaluates test with no availability guard (e.g. .enabled(if: SystemLanguageModel.default.isAvailable)). If the model is unavailable on the runner, every sample errors, every metric becomes .ignore, and the aggregate is computed over an empty set — which passes. A green gate that never ran is the worst outcome in this category.

Fix: pin the subject to greedy; put hard gates on pass/fail guardrails (which don't drift); size the scored-metric threshold above the measured noise floor. See axiom-ai (skills/foundation-models-evaluations.md) for the discipline and axiom-ai (skills/foundation-models-evaluations-diag.md) for the failure modes.

Category 2: Test Speed Issues (HIGH)

2.1 Host Application Not Needed

Detection: Unit tests with no UIKit/SwiftUI imports, no XCUIApplication usage Issue: Launching app adds 20-60 seconds per run Fix: Set Host Application to "None" for pure unit tests

2.2 Tests in App Target

Detection: Test files using @testable import MyApp that only test models/services/utilities Issue: App tests require simulator launch — 60x slower than package tests Fix: Extract testable logic into Swift Package, test with swift test

2.3 Unnecessary UI Test Overhead

Detection: Unit-style tests in UI test target Issue: UI tests have heavy setup/teardown Fix: Move to unit test target

Category 3: Swift Testing Migration (MEDIUM)

3.1 XCTestCase Migration Candidates

Search: XCTestCase with only basic XCTAssert* calls Issue: Missing modern testing features (parallelism, async, parameterization) Fix: Migrate to @Suite struct with @Test functions

3.2 Parameterized Test Opportunities

Detection: Multiple similar test functions (testParseValid, testParseInvalid, testParseEmpty) Issue: Repetitive tests that could be consolidated Fix: Use @Test(arguments:) parameterization

Category 4: Swift 6 Concurrency Issues (HIGH)

4.1 XCTestCase with MainActor Default

Search: class.*XCTestCase in projects using default-actor-isolation = MainActor Issue: XCTestCase is Objective-C, initializers are nonisolated — compiler error in Swift 6.2+ Fix:

// ❌ Error with MainActor default
final class MyTests: XCTestCase { }

// ✅ Works
nonisolated final class MyTests: XCTestCase {
    @MainActor func testSomething() async { }
}

4.2 Missing @MainActor on UI Tests

Detection: Tests accessing @MainActor types without isolation Issue: Swift 6 strict concurrency requires explicit isolation Fix: Add @MainActor to test function

Category 5: Test Quality Issues (MEDIUM/LOW)

5.1 Tests Without Assertions

Search: Test functions with no XCTAssert*, #expect, or #require Issue: Tests that don't assert don't verify behavior — false confidence Fix: Add meaningful assertions

5.2 Overly Long Setup

Detection: setUp() or setUpWithError() methods longer than 20 lines Issue: Complex setup makes tests hard to understand and maintain Fix: Extract to helper methods, use factory patterns

5.3 Force Unwrapping in Tests

Search: try!, as!, !. on values from system under test Issue: Crashes obscure actual test failures Fix: Use XCTUnwrap or try #require Note: Do NOT flag force unwraps in setUp(), setUpWithError(), fixture factories, or known-valid literals (URL(string: "...")!, UUID(uuidString: "...")!, NSRegularExpression(pattern: "...")!).

Phase 3: Reason About Test Completeness

Using the Coverage Shape Map from Phase 1 and your domain knowledge, check for what's untested — not just what's wrong with existing tests.

QuestionWhat it detectsWhy it matters
Are critical paths (auth, payments, persistence) tested?Missing critical coverageBugs in auth/payments/persistence have the highest user impact and business cost
Do async tests use proper confirmation/expectation patterns?Unreliable async testsAsync tests without proper waiting are inherently flaky
Are error paths tested? (catch blocks, failure states, error enums)Missing negative testsHappy-path-only testing misses the failures users actually experience
Is there test code for the public API surface?Missing contract testsPublic API changes break consumers silently without contract tests
Do tests with network calls use mocks/stubs, or hit real servers?Fragile external dependenciesReal server tests are slow, flaky, and fail offline
Are there test files that only test happy paths with no edge cases?Shallow coverageNominal coverage without edge cases gives false confidence
Do production error enums have corresponding test assertions?Untested error variantsEvery error case that can happen in production should be verified in tests

Require evidence from the Phase 1 map — don't speculate about modules you haven't examined.

Phase 4: Cross-Reference Findings

Bump severity for these combinations:

Finding A+ Finding B= CompoundSeverity
No tests for auth moduleAuth uses @MainActor + asyncUntested concurrency in security-critical codeCRITICAL
Missing error path teststry! in production codeCrash on unhandled errorCRITICAL
Test uses sleep()Tests auth flowFlaky test on critical pathCRITICAL
No tests for persistence layerDatabase migration code presentUntested migrations risk data lossHIGH
Tests exist but no assertions@testable import of payment moduleFalse confidence in payment codeHIGH
XCTestCase with shared mutable stateSwift 6 strict concurrency enabledData races in test infrastructureHIGH
No mock/stub for network layerTests import networking moduleFragile tests dependent on external serversMEDIUM

Also note overlaps with other auditors:

  • Untested @MainActor code → compound with concurrency auditor
  • Untested persistence migrations → compound with data auditor
  • Tests with sleep() in async context → compound with concurrency auditor

Phase 5: Test Health Score

## Test Health Score

| Metric | Value |
|--------|-------|
| Module coverage | X/Y production modules have tests (Z%) |
| Critical path coverage | auth (yes/no), payments (yes/no), persistence (yes/no), networking (yes/no) |
| Error path coverage | N error enums, M with test assertions (Z%) |
| Test reliability | N sleep() calls, M shared mutable state instances |
| Test speed | N tests requiring simulator, M pure unit tests |
| Test framework | N XCTest, M Swift Testing (migration %) |
| **Health** | **WELL TESTED / GAPS / UNDERTESTED** |

Scoring:

  • WELL TESTED: All critical paths tested, <3 flaky patterns, >70% module coverage, error paths covered
  • GAPS: Most critical paths tested, some flaky patterns or missing error coverage, or 40-70% module coverage
  • UNDERTESTED: Critical paths untested, or >5 flaky patterns, or <40% module coverage

Output Format

# Test Quality Audit Results

## Coverage Shape Map
[8-12 line summary from Phase 1]

## Summary
- CRITICAL: [N] issues
- HIGH: [N] issues
- MEDIUM: [N] issues
- LOW: [N] issues
- Phase 2 (anti-pattern detection): [N] issues
- Phase 3 (completeness reasoning): [N] issues
- Phase 4 (compound findings): [N] issues

## Test Health Score
[Phase 5 table]

## Issues by Severity

### [SEVERITY] [Category]: [Description]
**File**: path/to/file.swift:line (or module name for coverage gaps)
**Phase**: [2: Detection | 3: Completeness | 4: Compound]
**Issue**: What's wrong or missing
**Impact**: What happens if not fixed
**Fix**: Code example or recommended action
**Cross-Auditor Notes**: [if overlapping with another auditor]

## Quick Wins
1. [Fastest impact fix]
2. [Biggest speedup]
3. [Easiest migration]

## Recommendations
1. [Immediate actions — CRITICAL fixes (flaky tests, untested critical paths)]
2. [Short-term — HIGH fixes (speed improvements, Swift 6 compliance)]
3. [Long-term — coverage expansion from Phase 3 findings]

Output Limits

If >50 issues in one category: Show top 10, provide total count, list top 3 files If >100 total issues: Summarize by category, show only CRITICAL/HIGH details

False Positives (Not Issues)

  • sleep() in test helpers for rate limiting (check context)
  • static let constants (immutable is fine)
  • UI tests that legitimately need XCUIApplication
  • Performance tests using XCTMetric
  • Tests intentionally using XCTest for Objective-C interop
  • Force unwraps in setUp() / fixture setup on known-valid literals
  • Modules with no tests that are pure UI (better tested via UI tests or previews)

Related

For unit test patterns: axiom-testing (swift-testing reference) For UI test patterns: axiom-testing (ui-testing reference) For async test patterns: axiom-testing (testing-async reference) For flaky test diagnosis: test-failure-analyzer agent

Frequently asked questions about Testing Auditor Agent

Similar skills