Map the cloud. Scan it. Merge everything.
cloudg collects your infrastructure into a graph, fans out to the security scanners you already trust, and merges every finding against 28 compliance frameworks. It can also skip the scanning entirely: aggregate outputs you already have, or map the complete inventory, every service interlinked, with no scanner involved.
Installation
# core: graph engine, normaliser, renderers, ingest
pip install cloudg
# cloud SDK extras, pulled in only for collection
pip install cloudg[aws] # boto3, aioboto3
pip install cloudg[azure] # Azure management SDKs
pip install cloudg[gcp] # google-cloud-asset, google-auth
pip install cloudg[all]
The core package alone ships the graph engine, ontology builder, findings normaliser, report renderers and cloudg ingest. That is enough to aggregate existing scan results on a machine with no cloud access at all.
The scanners are not Python dependencies. Prowler, ScoutSuite, Checkov and Trivy are separate executables. Install whichever subset you want; cloudg checks PATH at run time and skips anything missing with a warning. install.sh / install.bat set up all of them, and the Docker image bundles the four binaries.
Requires Python 3.11 or newer.
The pipeline
collect -> graph -> scan -> normalise -> report
\-> ontology / RAG / terraform
Collection runs all configured providers concurrently with asyncio, iterating accounts and regions per provider. Regions auto-discover with --regions all. Assets and network edges go into a directed NetworkX graph: BFS from the internet node finds exposed resources, and blast-radius scoring estimates what an attacker could reach from each node.
The enabled scanners run in parallel over the same inventory, each in its own thread with its own timeout. Their findings, the graph's reachability findings and the built-in IAM linter's results all flow into the normaliser, which deduplicates within and across scanners, rescores against CVSS, and maps everything to compliance controls.
Every phase degrades gracefully. A missing binary, an unreachable provider or a failed exporter logs a warning; nothing else stops.
Alongside the pipeline sits an independent function: inventory mapping. It runs no scanners at all. It deep-collects everything deployed or default across every service, links it into one map, and lets scanner findings merge into that map later.
cloudg run
The full pipeline: collect, graph, scan, normalise, report.
cloudg run -p aws --regions us-east-1
cloudg run -p all --regions all
cloudg run -p aws --scanners prowler,trivy
cloudg run -p aws --regions us-east-1 --terraform
| Flag | Meaning |
|---|---|
-p, --provider | aws, azure, gcp or all; repeatable |
--regions | all for auto-discovery, or a comma-separated list |
--scanners | subset of prowler,scoutsuite,checkov,trivy,iam; defaults to scanners.enabled |
--iac-dir | directory for the IaC scanners (Checkov, Trivy fs) |
--images | comma-separated container images for Trivy |
--ontology / --no-ontology | RDF export, on by default |
--rag-export / --no-rag-export | RAG chunks, on by default |
--terraform / --no-terraform | Terraform recreation, off by default |
-o, --output | output directory, ./reports by default |
When no IaC directory is configured but --terraform is on, the IaC scanners target the generated Terraform recreation of the live infrastructure. Checkov then audits your actual cloud, not whatever directory cloudg happens to run from. There is deliberately no fallback to .: a zero-finding scan of an unrelated directory reads like a clean bill of health, and cloudg refuses to produce one.
collect & scan
The phases alone. collect writes assets and edges without running a scanner; scan runs the selected scanners concurrently and writes raw-findings.json, no collection.
cloudg collect -p aws --profile prod --region eu-west-1
cloudg scan -p aws --scanners prowler,checkov --iac-dir ./terraform
cloudg map
Scanner-independent inventory mapping: what exists, and how is it wired together. What is wrong stays the scanners' business. No scanner runs and none needs to be installed; read credentials are enough.
cloudg map -p aws --regions all
cloudg map -p all --regions all
cloudg map -p aws --findings ./reports/raw-findings.json # overlay scanner output
Coverage comes in three layers:
| Layer | What it adds |
|---|---|
| Deep collectors | everything the standard collectors know, plus the network fabric: route tables, internet/NAT gateways, network interfaces, volumes and disks, Elastic and public IPs, NACLs, VPC peering, transit gateways, customer-managed IAM policies |
| Catch-all sweeps | the AWS Resource Groups Tagging API, Azure Resource Manager's full resources.list(), and GCP Cloud Asset Inventory each enumerate every resource in scope, so services without a dedicated collector still land on the map; a duplicate of a dedicated collector's asset is dropped and the richer one wins |
| Relationship linker | derives edges from asset metadata alone: instance → security group, subnet ⊃ database, route table → gateway, Lambda → IAM role, secret → KMS key, CloudFront → origin bucket, VM → NIC → NSG chains, VPC peering, GCP IAM bindings, and a generic pass resolving any ARN or resource ID one asset holds to another collected asset |
Outputs: inventory-map.json (assets, interconnections, summary), inventory-map.graphml and inventory-graph.json. With --findings, additionally asset-map.json (per-asset risk, riskiest first) and compliance-map.json (framework → affected assets). The map never depends on the scanners; findings merge in whenever they exist.
The same capability is library API: engine.map_inventory(), or cloudg.inventory.InventoryMapper and RelationshipLinker standalone. See Python API.
cloudg ingest
Aggregation without execution. Feed cloudg the native output files of scans that already ran (in CI, on another host, on a schedule) and it runs the normalisation half of the pipeline: cross-scanner dedupe via the check-equivalence rulesets, compliance mapping and report generation. No scanner executes and no cloud credentials are needed.
cloudg ingest \
--prowler ./prowler-output/ \
--scoutsuite ./scoutsuite-report/ \
--checkov ./results_json.json \
--trivy ./trivy-image.json --trivy ./trivy-fs.json \
-o ./reports
Every input flag is repeatable and any combination or subset of tools works: one tool alone, or all four together. The exact formats are pinned down under Input requirements.
Ingest runs no collection, so outputs that need live assets (topology, ontology, RAG, Terraform) stay empty. With credentials available, combine both through the API: collect() + ingest_reports() + analyze().
cloudg report
Re-render reports from a previous run's findings.json. No cloud, no scanners. --format takes html, json, svg or all.
cloudg report -i ./reports/findings.json --format html
Input requirements
What each ingest path expects, and the command that produces it.
Prowler
prowler aws -M json-asff -o ./prowler-output
cloudg accepts the output directory or any single file from it, as a JSON array or JSONL. The fields read from each ASFF finding:
{
"Id": "prowler-aws-s3_bucket_default_encryption-123456789012-eu-west-1-...",
"Title": "S3 bucket default encryption",
"Severity": {"Label": "HIGH"},
"Resources": [{"Id": "arn:aws:s3:::my-bucket"}],
"Compliance": {"Status": "FAILED", "RelatedRequirements": ["CIS 2.1.1"]},
"Remediation": {"Recommendation": {"Text": "..."}}
}
Findings with a PASSED status are skipped. The check name embedded in Id becomes the dedupe key.
ScoutSuite
scout aws --report-dir ./scoutsuite-report --no-browser
Accepts the report directory (searched recursively for scoutsuite_results*.js) or the file itself. The file is JavaScript, a JSON object assigned to a variable; cloudg strips the assignment and parses the JSON. Each entry under services.<service>.findings with flagged_items > 0 yields one finding per flagged item. level maps to severity: danger is CRITICAL, warning HIGH, caution MEDIUM.
Checkov
checkov -d ./iac --output json > results_json.json
Accepts the JSON file or a directory containing results_json.json. Both output shapes parse: a single {check_type, results} object, or a list of them when several frameworks ran. Findings come from results.failed_checks[], reading check_id, check_name, severity, resource, file_path and guideline.
Trivy
trivy image --format json myrepo/app:latest > trivy-image.json
trivy fs --format json --scanners vuln,misconfig,secret ./iac > trivy-fs.json
Accepts a single JSON file or a directory of them. The top-level ArtifactType decides the parser: container_image goes through the image parser (vulnerabilities and secrets), everything else through the filesystem parser, which also handles misconfigurations. ArtifactName becomes the resource identifier, and the CVSS v3 score is read when present.
Output contract
| File | Contents |
|---|---|
report.html | interactive report: D3 topology, findings table, compliance matrix; self-contained, works offline |
findings.json | the machine-readable result, structure below; round-trips through cloudg report -i |
raw-findings.json | pre-normalisation findings from scan and ingest |
topology.svg / .graphml / -cytoscape.json | the graph in three formats |
ontology.ttl / .jsonld | RDF ontology, around 62 inferred relation types, SPARQL-queryable |
rag_chunks.jsonl | retrieval-ready chunks, one JSON object per line |
terraform/*.tf.json | Terraform recreation of live infrastructure with an import.sh |
inventory-map.json / .graphml, inventory-graph.json | scanner-independent inventory map: assets, interconnections, summary (cloudg map) |
asset-map.json, compliance-map.json | inventory overlaid with scanner findings (cloudg map --findings) |
{
"metadata": {"scan_id", "provider", "account_id", "region", "started_at", "completed_at"},
"summary": {"total_assets", "total_findings", "severity_breakdown", "compliance_frameworks"},
"assets": [CloudAsset, ...],
"findings": [Finding, ...],
"compliance": [ComplianceResult, ...],
"graph": {"nodes": [], "links": []}
}
Data models
Everything flowing through cloudg is a Pydantic v2 model from cloudg.schema.models. All models allow extra fields, so scanner-specific attributes survive serialisation.
Finding
| Field | Type | Notes |
|---|---|---|
resource_id, resource_arn | str | internal ID and cloud-native identifier |
severity | Severity | CRITICAL, HIGH, MEDIUM, LOW, INFO |
title, description | str | scanner tags are stripped for cross-scanner comparison |
evidence, remediation | str | None | |
source_tool | str | "prowler"; becomes "prowler, checkov" after a merge |
source_finding_id | str | None | the scanner's own check or CVE ID; drives dedupe |
compliance_frameworks | list[str] | unioned on merge |
cvss_score | float | None | 0.0 to 10.0 |
risk_score | computed | severity base averaged with CVSS when present |
The rest
CloudAsset carries arn, name, a 45-value asset_type taxonomy (EC2, S3_BUCKET, IAM_ROLE, VPC, KMS_KEY, ...), provider, region, tags, metadata and is_internet_exposed. NetworkEdge is a directed edge with an edge_type (SECURITY_GROUP_RULE, IAM_TRUST, CONTAINS, INTERNET_EXPOSED, ATTACHED_TO, REFERENCES, ...), ports, protocol and CIDR. ComplianceResult is one control-level verdict with framework, control_id, status and the finding_ids behind it. ScanResult is the aggregate container the normaliser returns and the renderers consume, with a computed summary.
Configuration
config.yaml mirrors every CLI flag and adds persistent settings. Everything is optional. The shipped config.yaml documents each field inline; the structure, with defaults:
providers: [aws]
aws:
regions: [us-east-1] # or [all]
accounts: [] # multi-account fan-out
role_name: null
role_arn: null
external_id: null
web_identity_token_file: null
scanners:
enabled: [prowler, scoutsuite, checkov, trivy, iam]
timeout_seconds: 3600
iac_directories: []
trivy_images: []
prowler_extra_args: [] # passthrough, per scanner
inventory: { tagging_sweep: true, link_references: true } # cloudg map
ontology: { enabled: true, export_formats: [turtle, json-ld] }
rag: { enabled: true, chunk_strategy: hybrid, max_chunk_tokens: 2000 }
terraform: { enabled: false }
rulesets: { rules_dir: null } # defaults to the packaged cloudg/rules/
concurrency_limit: 5
In code, the same structure is CloudGConfig:
from cloudg import CloudGConfig, load_config
config = load_config("config.yaml")
config = CloudGConfig(providers=["aws"])
config.scanners.enabled = ["prowler", "iam"]
Python API
CloudGEngine in cloudg.api is the integration entry point: async-first, sync wrappers included, built for embedding in larger systems (SIEM pipelines, orchestration platforms, command extensions such as hol-guard).
from cloudg import CloudGConfig, CloudGEngine
engine = CloudGEngine(CloudGConfig(providers=["aws"]))
engine.on_finding = lambda f: siem.send(f.model_dump(mode="json"))
engine.on_error = lambda phase, exc: alerting.notify(phase, exc)
result = engine.run_pipeline_sync()
print(result.to_summary())
| Method | Returns | Does |
|---|---|---|
collect() | CollectionResult | multi-provider asset collection: assets, edges, coverage |
scan(assets, edges) | list[Finding] | enabled scanners in parallel, plus reachability and IAM linting |
ingest_reports(reports) | list[Finding] | parse existing scanner outputs; nothing executes |
normalise_findings(findings) | ScanResult | dedupe, cross-scanner merge, CVSS rescoring, compliance mapping |
analyze(assets, edges, findings) | AnalysisResult | graph, ontology, RAG chunks, Terraform, attack paths |
run_pipeline() | PipelineResult | everything, end to end; run_pipeline_sync() for non-async callers |
run_from_reports(reports) | PipelineResult | ingest, normalise, report in one call; sync twin available |
map_inventory() | InventoryResult | scanner-independent inventory mapping: deep collection, catch-all sweeps, relationship linking; sync twin available |
Event hooks: on_phase_start(phase), on_collection_complete(result), on_finding(finding), on_scan_complete(findings), on_analysis_complete(result), on_error(phase, exc). Exceptions raised inside a hook are swallowed, so a broken callback cannot take the pipeline down. PipelineResult carries the data, report paths, coverage and a to_summary() dict ready to serialise.
Inventory mapping from code
from cloudg.inventory import InventoryMapper, RelationshipLinker
mapper = InventoryMapper(config)
inventory = mapper.map_inventory_sync() # no scanners
inventory.export("./reports") # map + graphml + graph json
asset_map = mapper.build_asset_map(inventory, findings) # asset -> risk
compliance = mapper.build_compliance_map(inventory, findings)
edges = RelationshipLinker(assets).link() # works on ANY CloudAsset list
The deeper toolkit
The CLI is a thin layer; the modules the pipeline uses internally are public, importable API. The markdown reference documents each in full.
| Module | From code |
|---|---|
cloudg.inventory | InventoryMapper, RelationshipLinker, and the three deep collectors, all standalone |
cloudg.graph.builder | GraphBuilder: attack paths, lateral movement, centrality / blast-radius metrics, D3 / Cytoscape / GraphML export |
cloudg.graph.ontology | CloudOntology: ~62 typed RDF relations, SPARQL-queryable, Turtle / JSON-LD |
cloudg.graph.rag_export | RAGExporter: retrieval-ready JSONL chunks of the infrastructure for LLM pipelines |
cloudg.renderers.terraform_export | TerraformExporter: .tf.json recreation plus import.sh, with a coverage preview() |
cloudg.normaliser | FindingsNormaliser: dedupe, cross-scanner merge, CVSS rescore, compliance mapping on any findings list |
cloudg.ingest | parse_report / ingest_reports: every scanner's parser, no engine, no credentials |
Integration recipes
Aggregate existing scan results, get reports
engine = CloudGEngine(CloudGConfig())
result = engine.run_from_reports_sync(
{"prowler": ["./prowler-output/"], "trivy": ["./trivy-image.json"]},
output_dir="./reports",
)
print(result.total_findings, result.severity_breakdown)
print(result.report_paths["html"])
Parse first, decide later
from cloudg.ingest import ingest_reports
findings = ingest_reports({"prowler": ["./out/"], "checkov": ["results_json.json"]})
scan_result = engine.normalise_findings(findings)
for c in scan_result.compliance:
if c.status.value == "FAIL":
print(c.framework, c.control_id, len(c.finding_ids))
The per-scanner parsers are also public: TrivyScanner.parse_report("./trivy.json") and friends, one tool with no engine.
Map the inventory, overlay findings later
inventory = engine.map_inventory_sync(output_dir="./reports") # no scanners
print(inventory.summary["assets_by_service"]) # what exists
print(inventory.summary["internet_exposed"]) # what faces the internet
# hours or days later, when scanner output exists:
findings = engine.ingest_reports({"prowler": ["./ci-prowler-output/"]})
InventoryMapper(config).export_merged(inventory, findings, "./reports")
# -> asset-map.json, compliance-map.json
Combine live collection with ingested reports
collection = await engine.collect()
ingested = engine.ingest_reports({"prowler": ["./ci-prowler-output/"]})
scan_result = engine.normalise_findings(ingested, assets=collection.assets)
await engine.analyze(collection.assets, collection.edges, scan_result.findings)
Gate a CI job on severity
result = CloudGEngine(CloudGConfig()).run_from_reports_sync(
{"checkov": ["results_json.json"]}
)
if result.severity_breakdown.get("CRITICAL", 0):
sys.exit(1)
Plugins
Custom collectors and scanners register through setuptools entry points; no core changes needed. PluginRegistry discovers them at startup and falls back to the built-ins for any name not overridden.
# my_package/scanner.py
from cloudg.schema.models import Finding, Severity
class MyScanner:
def run(self) -> list[Finding]:
return [Finding(
resource_id="arn:aws:s3:::example",
severity=Severity.HIGH,
title="Example finding",
description="What was found and why it matters",
source_tool="myscanner",
source_finding_id="MY_CHECK_001",
)]
[project.entry-points."cloudg.scanners"]
myscanner = "my_package.scanner:MyScanner"
[project.entry-points."cloudg.collectors"]
mycloud = "my_package.collector:MyCollector"
After pip install, the name works in scanners.enabled and --scanners. If your scanner emits a stable check ID in source_finding_id, add it to rules/check_equivalence.yaml so its findings merge with equivalent checks from other tools.
Authentication
Every provider supports several methods, resolved in fixed priority order, so the same config works on a laptop, in CI and on cloud compute.
| Provider | Methods, most explicit first |
|---|---|
| AWS | direct keys; OIDC web identity federation (--aws-role-arn with --aws-web-identity-token-file, the GitHub Actions and EKS pattern); named profile including SSO; the default chain. STS assumption layers on top with --aws-external-id for the third-party auditor pattern, and multi-account fan-out uses accounts plus role_name. |
| Azure | workload identity federation; service principal with secret or certificate; managed identity; the DefaultAzureCredential chain, which covers az login. |
| GCP | credentials file (service account key or workload identity federation config); application default credentials. --gcp-impersonate-sa layers impersonation on either. |
Docker
docker pull ghcr.io/morpheuslord/cloudg:latest
docker compose run --rm cloudg run -p aws --regions us-east-1
The image bundles all four scanner binaries, so it is the shortest path to the full pipeline.
Changelog
Major releases only. The full log lives in CHANGELOG.md.
Inventory mapping
A scanner-independent inventory mapper: cloudg map and CloudGEngine.map_inventory() deep-collect everything deployed or default across AWS, Azure and GCP (network fabric included, with catch-all sweeps so every service appears) and link it into one map. Scanner findings merge in afterwards as asset-map.json and compliance-map.json. The relationship linker, deep collectors and mapper are all public API.
Ingest mode
cloudg now works entirely from scanner outputs you already have: the cloudg ingest command, public parse_report() parsers on all four scanner wrappers, and ingest_reports() / normalise_findings() / run_from_reports() on the engine. No cloud credentials, no scanner binaries. Plus this documentation and a least-privilege fix to the publish workflow.
The cloudg rebrand
The project, package, CLI, Docker image, installers and CI became cloudg. Multi-cloud authentication (OIDC federation, STS assumption, managed identity, impersonation), compliance rulesets generated from Prowler's public data (28 frameworks, 4,166 controls, 10,236 check mappings), uv packaging, PyPI trusted publishing, GHCR image, MIT license. Patches folded in: 0.3.1 stopped the IaC scanners from silently scanning the working directory; 0.3.2 rekeyed dedupe on scanner check IDs with the cross-scanner equivalence map and moved the CLI to a Rich UI layer.