AI-Assisted Penetration Testing: A Practical Evaluation Guide
Published February 4, 2025 · Updated October 6, 2026 · David Ige

Reviewed October 6, 2026. AI-assisted penetration testing should be evaluated as a workflow with an operator in control, not by unsupported claims that a tool is more capable or saves a fixed amount of time. This guide gives a small, repeatable way to assess Nebula against sanitized assessment records.
Version status: stable and preview are different channels
As checked on October 6, 2026, the signed Nebula APT channel lists no stable package and one published prerelease: Nebula 3.0.0-alpha.5. Treat that build as preview software. The source repository can contain notes or code for later unpublished candidates; a source version is not proof that a release is available. Pin the published version, platform, and model provider for every evaluation, and recheck the release page and signed channel manifest before installing.
What current Nebula does
Nebula 3 is documented as a security workbench for authorized assessments: it brings terminal, notes, findings, reports, and evidence into one workflow. It supports local, hosted, and OpenAI-compatible model runtimes, and the model provider is optional. Those facts replace the old article’s claims about a fixed list of models downloaded at first run and a single CLI interaction pattern.
Local and hosted models
A local model sends inference to a local runtime, subject to the configured endpoint and its integrations. A hosted model sends the selected prompt and context to that provider. Search, browser, MCP, or other integrations can also transmit data independently of model location. Record the provider, endpoint, model ID, and enabled integrations; use only sanitized assessment data unless the data owner has approved the path.
Scope, approvals, and evidence
Set the authorized target scope and policy before testing. Nebula documents approval pauses, execution budgets, and isolated container execution. An approval authorizes a proposed action within that configured boundary; it does not prove that the action is safe or that an AI-generated finding is correct. Review commands and outputs before use.
Carry claims into a report only when they point to supporting assessment evidence. Preserve artifact identifiers, source output, execution provenance, and the exact Nebula version. Inspect exports and diagnostic bundles for sensitive data before sharing them.
A small evaluation with sanitized records
Use eight records: four with independently confirmed findings and four benign or ambiguous cases. Remove names, credentials, real hostnames, IP addresses, and client-identifying text. Keep a human-reviewed reference answer and source evidence for each record. Freeze the Nebula version, provider, model, prompt, and record set; compare a human-only report with the AI-assisted draft under the same review standard.
| Measure | How to score it | Result in this revision |
|---|---|---|
| Summary accuracy | Fact-check each material claim against the reference; report supported claims and material omissions | Not measured |
| Evidence traceability | Share of reported findings with an exact record/artifact reference and source location | Not measured |
| Invented findings | Count findings not supported by the sanitized record; report the rate and severity separately | Not measured |
| Time to reviewed report | Elapsed time from opening the record to a human-approved report, including corrections | Not measured |
No sanitized record set or run artifacts were available for this revision, so the table contains no performance result. Do not claim better accuracy, superiority, or time saved until this evaluation has been run and reviewed. Record model calls, reviewer corrections, and failed or refused actions as part of the evidence.
Where Nebula ends and DAP begins
Nebula supports the authorized assessment workflow: scoped testing, operator-reviewed execution, notes, findings, and reports tied to evidence. DAP is a separate, cloud-based binary-analysis workflow for supported executable formats, including PE, ELF, and managed .NET files. It does not perform Nebula’s network assessment or validate Nebula’s findings. DAP’s analysis backend is hosted, not part of Nebula’s local model runtime; submit a sample only when its handling is allowed by the assessment policy.
For the separate task of reviewing a suspicious executable, see DAP’s binary-analysis workflow.
DeepExploit attribution
DeepExploit is not a 2019 MIT prototype. Its original public project identifies Isao Takaesu (@13o-bbr-bbq) as its developer and describes a machine-learning penetration-testing tool. It is a separate project from Nebula. See the original DeepExploit source repository.