cloud graphing · AWS / Azure / GCP

Map the cloud. Scan it. Merge everything.

cloudg collects your infrastructure into a graph, fans out to the security scanners you already trust, and merges every finding against 28 compliance frameworks. It can also skip the scanning entirely: aggregate outputs you already have, or map the complete inventory, every service interlinked, with no scanner involved.

Installation

# core: graph engine, normaliser, renderers, ingest
pip install cloudg

# cloud SDK extras, pulled in only for collection
pip install cloudg[aws]        # boto3, aioboto3
pip install cloudg[azure]      # Azure management SDKs
pip install cloudg[gcp]        # google-cloud-asset, google-auth
pip install cloudg[all]

The core package alone ships the graph engine, ontology builder, findings normaliser, report renderers and cloudg ingest. That is enough to aggregate existing scan results on a machine with no cloud access at all.

The scanners are not Python dependencies. Prowler, ScoutSuite, Checkov and Trivy are separate executables. Install whichever subset you want; cloudg checks PATH at run time and skips anything missing with a warning. install.sh / install.bat set up all of them, and the Docker image bundles the four binaries.

Requires Python 3.11 or newer.

The pipeline

collect -> graph -> scan -> normalise -> report
                \-> ontology / RAG / terraform

Collection runs all configured providers concurrently with asyncio, iterating accounts and regions per provider. Regions auto-discover with --regions all. Assets and network edges go into a directed NetworkX graph: BFS from the internet node finds exposed resources, and blast-radius scoring estimates what an attacker could reach from each node.

The enabled scanners run in parallel over the same inventory, each in its own thread with its own timeout. Their findings, the graph's reachability findings and the built-in IAM linter's results all flow into the normaliser, which deduplicates within and across scanners, rescores against CVSS, and maps everything to compliance controls.

Every phase degrades gracefully. A missing binary, an unreachable provider or a failed exporter logs a warning; nothing else stops.

Alongside the pipeline sits an independent function: inventory mapping. It runs no scanners at all. It deep-collects everything deployed or default across every service, links it into one map, and lets scanner findings merge into that map later.

cloudg run

The full pipeline: collect, graph, scan, normalise, report.

cloudg run -p aws --regions us-east-1
cloudg run -p all --regions all
cloudg run -p aws --scanners prowler,trivy
cloudg run -p aws --regions us-east-1 --terraform
FlagMeaning
-p, --provideraws, azure, gcp or all; repeatable
--regionsall for auto-discovery, or a comma-separated list
--scannerssubset of prowler,scoutsuite,checkov,trivy,iam; defaults to scanners.enabled
--iac-dirdirectory for the IaC scanners (Checkov, Trivy fs)
--imagescomma-separated container images for Trivy
--ontology / --no-ontologyRDF export, on by default
--rag-export / --no-rag-exportRAG chunks, on by default
--terraform / --no-terraformTerraform recreation, off by default
-o, --outputoutput directory, ./reports by default

When no IaC directory is configured but --terraform is on, the IaC scanners target the generated Terraform recreation of the live infrastructure. Checkov then audits your actual cloud, not whatever directory cloudg happens to run from. There is deliberately no fallback to .: a zero-finding scan of an unrelated directory reads like a clean bill of health, and cloudg refuses to produce one.

collect & scan

The phases alone. collect writes assets and edges without running a scanner; scan runs the selected scanners concurrently and writes raw-findings.json, no collection.

cloudg collect -p aws --profile prod --region eu-west-1
cloudg scan -p aws --scanners prowler,checkov --iac-dir ./terraform

cloudg map

Scanner-independent inventory mapping: what exists, and how is it wired together. What is wrong stays the scanners' business. No scanner runs and none needs to be installed; read credentials are enough.

cloudg map -p aws --regions all
cloudg map -p all --regions all
cloudg map -p aws --findings ./reports/raw-findings.json   # overlay scanner output

Coverage comes in three layers:

LayerWhat it adds
Deep collectorseverything the standard collectors know, plus the network fabric: route tables, internet/NAT gateways, network interfaces, volumes and disks, Elastic and public IPs, NACLs, VPC peering, transit gateways, customer-managed IAM policies
Catch-all sweepsthe AWS Resource Groups Tagging API, Azure Resource Manager's full resources.list(), and GCP Cloud Asset Inventory each enumerate every resource in scope, so services without a dedicated collector still land on the map; a duplicate of a dedicated collector's asset is dropped and the richer one wins
Relationship linkerderives edges from asset metadata alone: instance → security group, subnet ⊃ database, route table → gateway, Lambda → IAM role, secret → KMS key, CloudFront → origin bucket, VM → NIC → NSG chains, VPC peering, GCP IAM bindings, and a generic pass resolving any ARN or resource ID one asset holds to another collected asset

Outputs: inventory-map.json (assets, interconnections, summary), inventory-map.graphml and inventory-graph.json. With --findings, additionally asset-map.json (per-asset risk, riskiest first) and compliance-map.json (framework → affected assets). The map never depends on the scanners; findings merge in whenever they exist.

The same capability is library API: engine.map_inventory(), or cloudg.inventory.InventoryMapper and RelationshipLinker standalone. See Python API.

cloudg ingest

Aggregation without execution. Feed cloudg the native output files of scans that already ran (in CI, on another host, on a schedule) and it runs the normalisation half of the pipeline: cross-scanner dedupe via the check-equivalence rulesets, compliance mapping and report generation. No scanner executes and no cloud credentials are needed.

cloudg ingest \
  --prowler ./prowler-output/ \
  --scoutsuite ./scoutsuite-report/ \
  --checkov ./results_json.json \
  --trivy ./trivy-image.json --trivy ./trivy-fs.json \
  -o ./reports

Every input flag is repeatable and any combination or subset of tools works: one tool alone, or all four together. The exact formats are pinned down under Input requirements.

Ingest runs no collection, so outputs that need live assets (topology, ontology, RAG, Terraform) stay empty. With credentials available, combine both through the API: collect() + ingest_reports() + analyze().

cloudg report

Re-render reports from a previous run's findings.json. No cloud, no scanners. --format takes html, json, svg or all.

cloudg report -i ./reports/findings.json --format html

Input requirements

What each ingest path expects, and the command that produces it.

Prowler

prowler aws -M json-asff -o ./prowler-output

cloudg accepts the output directory or any single file from it, as a JSON array or JSONL. The fields read from each ASFF finding:

{
  "Id": "prowler-aws-s3_bucket_default_encryption-123456789012-eu-west-1-...",
  "Title": "S3 bucket default encryption",
  "Severity": {"Label": "HIGH"},
  "Resources": [{"Id": "arn:aws:s3:::my-bucket"}],
  "Compliance": {"Status": "FAILED", "RelatedRequirements": ["CIS 2.1.1"]},
  "Remediation": {"Recommendation": {"Text": "..."}}
}

Findings with a PASSED status are skipped. The check name embedded in Id becomes the dedupe key.

ScoutSuite

scout aws --report-dir ./scoutsuite-report --no-browser

Accepts the report directory (searched recursively for scoutsuite_results*.js) or the file itself. The file is JavaScript, a JSON object assigned to a variable; cloudg strips the assignment and parses the JSON. Each entry under services.<service>.findings with flagged_items > 0 yields one finding per flagged item. level maps to severity: danger is CRITICAL, warning HIGH, caution MEDIUM.

Checkov

checkov -d ./iac --output json > results_json.json

Accepts the JSON file or a directory containing results_json.json. Both output shapes parse: a single {check_type, results} object, or a list of them when several frameworks ran. Findings come from results.failed_checks[], reading check_id, check_name, severity, resource, file_path and guideline.

Trivy

trivy image --format json myrepo/app:latest > trivy-image.json
trivy fs --format json --scanners vuln,misconfig,secret ./iac > trivy-fs.json

Accepts a single JSON file or a directory of them. The top-level ArtifactType decides the parser: container_image goes through the image parser (vulnerabilities and secrets), everything else through the filesystem parser, which also handles misconfigurations. ArtifactName becomes the resource identifier, and the CVSS v3 score is read when present.

Output contract

FileContents
report.htmlinteractive report: D3 topology, findings table, compliance matrix; self-contained, works offline
findings.jsonthe machine-readable result, structure below; round-trips through cloudg report -i
raw-findings.jsonpre-normalisation findings from scan and ingest
topology.svg / .graphml / -cytoscape.jsonthe graph in three formats
ontology.ttl / .jsonldRDF ontology, around 62 inferred relation types, SPARQL-queryable
rag_chunks.jsonlretrieval-ready chunks, one JSON object per line
terraform/*.tf.jsonTerraform recreation of live infrastructure with an import.sh
inventory-map.json / .graphml, inventory-graph.jsonscanner-independent inventory map: assets, interconnections, summary (cloudg map)
asset-map.json, compliance-map.jsoninventory overlaid with scanner findings (cloudg map --findings)
{
  "metadata":   {"scan_id", "provider", "account_id", "region", "started_at", "completed_at"},
  "summary":    {"total_assets", "total_findings", "severity_breakdown", "compliance_frameworks"},
  "assets":     [CloudAsset, ...],
  "findings":   [Finding, ...],
  "compliance": [ComplianceResult, ...],
  "graph":      {"nodes": [], "links": []}
}

Data models

Everything flowing through cloudg is a Pydantic v2 model from cloudg.schema.models. All models allow extra fields, so scanner-specific attributes survive serialisation.

Finding

FieldTypeNotes
resource_id, resource_arnstrinternal ID and cloud-native identifier
severitySeverityCRITICAL, HIGH, MEDIUM, LOW, INFO
title, descriptionstrscanner tags are stripped for cross-scanner comparison
evidence, remediationstr | None
source_toolstr"prowler"; becomes "prowler, checkov" after a merge
source_finding_idstr | Nonethe scanner's own check or CVE ID; drives dedupe
compliance_frameworkslist[str]unioned on merge
cvss_scorefloat | None0.0 to 10.0
risk_scorecomputedseverity base averaged with CVSS when present

The rest

CloudAsset carries arn, name, a 45-value asset_type taxonomy (EC2, S3_BUCKET, IAM_ROLE, VPC, KMS_KEY, ...), provider, region, tags, metadata and is_internet_exposed. NetworkEdge is a directed edge with an edge_type (SECURITY_GROUP_RULE, IAM_TRUST, CONTAINS, INTERNET_EXPOSED, ATTACHED_TO, REFERENCES, ...), ports, protocol and CIDR. ComplianceResult is one control-level verdict with framework, control_id, status and the finding_ids behind it. ScanResult is the aggregate container the normaliser returns and the renderers consume, with a computed summary.

Configuration

config.yaml mirrors every CLI flag and adds persistent settings. Everything is optional. The shipped config.yaml documents each field inline; the structure, with defaults:

providers: [aws]

aws:
  regions: [us-east-1]          # or [all]
  accounts: []                  # multi-account fan-out
  role_name: null
  role_arn: null
  external_id: null
  web_identity_token_file: null

scanners:
  enabled: [prowler, scoutsuite, checkov, trivy, iam]
  timeout_seconds: 3600
  iac_directories: []
  trivy_images: []
  prowler_extra_args: []        # passthrough, per scanner

inventory: { tagging_sweep: true, link_references: true }  # cloudg map
ontology:  { enabled: true, export_formats: [turtle, json-ld] }
rag:       { enabled: true, chunk_strategy: hybrid, max_chunk_tokens: 2000 }
terraform: { enabled: false }
rulesets:  { rules_dir: null }  # defaults to the packaged cloudg/rules/
concurrency_limit: 5

In code, the same structure is CloudGConfig:

from cloudg import CloudGConfig, load_config

config = load_config("config.yaml")
config = CloudGConfig(providers=["aws"])
config.scanners.enabled = ["prowler", "iam"]

Python API

CloudGEngine in cloudg.api is the integration entry point: async-first, sync wrappers included, built for embedding in larger systems (SIEM pipelines, orchestration platforms, command extensions such as hol-guard).

from cloudg import CloudGConfig, CloudGEngine

engine = CloudGEngine(CloudGConfig(providers=["aws"]))
engine.on_finding = lambda f: siem.send(f.model_dump(mode="json"))
engine.on_error = lambda phase, exc: alerting.notify(phase, exc)

result = engine.run_pipeline_sync()
print(result.to_summary())
MethodReturnsDoes
collect()CollectionResultmulti-provider asset collection: assets, edges, coverage
scan(assets, edges)list[Finding]enabled scanners in parallel, plus reachability and IAM linting
ingest_reports(reports)list[Finding]parse existing scanner outputs; nothing executes
normalise_findings(findings)ScanResultdedupe, cross-scanner merge, CVSS rescoring, compliance mapping
analyze(assets, edges, findings)AnalysisResultgraph, ontology, RAG chunks, Terraform, attack paths
run_pipeline()PipelineResulteverything, end to end; run_pipeline_sync() for non-async callers
run_from_reports(reports)PipelineResultingest, normalise, report in one call; sync twin available
map_inventory()InventoryResultscanner-independent inventory mapping: deep collection, catch-all sweeps, relationship linking; sync twin available

Event hooks: on_phase_start(phase), on_collection_complete(result), on_finding(finding), on_scan_complete(findings), on_analysis_complete(result), on_error(phase, exc). Exceptions raised inside a hook are swallowed, so a broken callback cannot take the pipeline down. PipelineResult carries the data, report paths, coverage and a to_summary() dict ready to serialise.

Inventory mapping from code

from cloudg.inventory import InventoryMapper, RelationshipLinker

mapper = InventoryMapper(config)
inventory = mapper.map_inventory_sync()                 # no scanners
inventory.export("./reports")                          # map + graphml + graph json

asset_map = mapper.build_asset_map(inventory, findings) # asset -> risk
compliance = mapper.build_compliance_map(inventory, findings)

edges = RelationshipLinker(assets).link()               # works on ANY CloudAsset list

The deeper toolkit

The CLI is a thin layer; the modules the pipeline uses internally are public, importable API. The markdown reference documents each in full.

ModuleFrom code
cloudg.inventoryInventoryMapper, RelationshipLinker, and the three deep collectors, all standalone
cloudg.graph.builderGraphBuilder: attack paths, lateral movement, centrality / blast-radius metrics, D3 / Cytoscape / GraphML export
cloudg.graph.ontologyCloudOntology: ~62 typed RDF relations, SPARQL-queryable, Turtle / JSON-LD
cloudg.graph.rag_exportRAGExporter: retrieval-ready JSONL chunks of the infrastructure for LLM pipelines
cloudg.renderers.terraform_exportTerraformExporter: .tf.json recreation plus import.sh, with a coverage preview()
cloudg.normaliserFindingsNormaliser: dedupe, cross-scanner merge, CVSS rescore, compliance mapping on any findings list
cloudg.ingestparse_report / ingest_reports: every scanner's parser, no engine, no credentials

Integration recipes

Aggregate existing scan results, get reports

engine = CloudGEngine(CloudGConfig())
result = engine.run_from_reports_sync(
    {"prowler": ["./prowler-output/"], "trivy": ["./trivy-image.json"]},
    output_dir="./reports",
)
print(result.total_findings, result.severity_breakdown)
print(result.report_paths["html"])

Parse first, decide later

from cloudg.ingest import ingest_reports

findings = ingest_reports({"prowler": ["./out/"], "checkov": ["results_json.json"]})
scan_result = engine.normalise_findings(findings)

for c in scan_result.compliance:
    if c.status.value == "FAIL":
        print(c.framework, c.control_id, len(c.finding_ids))

The per-scanner parsers are also public: TrivyScanner.parse_report("./trivy.json") and friends, one tool with no engine.

Map the inventory, overlay findings later

inventory = engine.map_inventory_sync(output_dir="./reports")   # no scanners
print(inventory.summary["assets_by_service"])   # what exists
print(inventory.summary["internet_exposed"])    # what faces the internet

# hours or days later, when scanner output exists:
findings = engine.ingest_reports({"prowler": ["./ci-prowler-output/"]})
InventoryMapper(config).export_merged(inventory, findings, "./reports")
# -> asset-map.json, compliance-map.json

Combine live collection with ingested reports

collection = await engine.collect()
ingested = engine.ingest_reports({"prowler": ["./ci-prowler-output/"]})
scan_result = engine.normalise_findings(ingested, assets=collection.assets)
await engine.analyze(collection.assets, collection.edges, scan_result.findings)

Gate a CI job on severity

result = CloudGEngine(CloudGConfig()).run_from_reports_sync(
    {"checkov": ["results_json.json"]}
)
if result.severity_breakdown.get("CRITICAL", 0):
    sys.exit(1)

Plugins

Custom collectors and scanners register through setuptools entry points; no core changes needed. PluginRegistry discovers them at startup and falls back to the built-ins for any name not overridden.

# my_package/scanner.py
from cloudg.schema.models import Finding, Severity

class MyScanner:
    def run(self) -> list[Finding]:
        return [Finding(
            resource_id="arn:aws:s3:::example",
            severity=Severity.HIGH,
            title="Example finding",
            description="What was found and why it matters",
            source_tool="myscanner",
            source_finding_id="MY_CHECK_001",
        )]
[project.entry-points."cloudg.scanners"]
myscanner = "my_package.scanner:MyScanner"

[project.entry-points."cloudg.collectors"]
mycloud = "my_package.collector:MyCollector"

After pip install, the name works in scanners.enabled and --scanners. If your scanner emits a stable check ID in source_finding_id, add it to rules/check_equivalence.yaml so its findings merge with equivalent checks from other tools.

Authentication

Every provider supports several methods, resolved in fixed priority order, so the same config works on a laptop, in CI and on cloud compute.

ProviderMethods, most explicit first
AWSdirect keys; OIDC web identity federation (--aws-role-arn with --aws-web-identity-token-file, the GitHub Actions and EKS pattern); named profile including SSO; the default chain. STS assumption layers on top with --aws-external-id for the third-party auditor pattern, and multi-account fan-out uses accounts plus role_name.
Azureworkload identity federation; service principal with secret or certificate; managed identity; the DefaultAzureCredential chain, which covers az login.
GCPcredentials file (service account key or workload identity federation config); application default credentials. --gcp-impersonate-sa layers impersonation on either.

Docker

docker pull ghcr.io/morpheuslord/cloudg:latest
docker compose run --rm cloudg run -p aws --regions us-east-1

The image bundles all four scanner binaries, so it is the shortest path to the full pipeline.

Changelog

Major releases only. The full log lives in CHANGELOG.md.

0.5.0 unreleased

Inventory mapping

A scanner-independent inventory mapper: cloudg map and CloudGEngine.map_inventory() deep-collect everything deployed or default across AWS, Azure and GCP (network fabric included, with catch-all sweeps so every service appears) and link it into one map. Scanner findings merge in afterwards as asset-map.json and compliance-map.json. The relationship linker, deep collectors and mapper are all public API.

0.4.0 unreleased

Ingest mode

cloudg now works entirely from scanner outputs you already have: the cloudg ingest command, public parse_report() parsers on all four scanner wrappers, and ingest_reports() / normalise_findings() / run_from_reports() on the engine. No cloud credentials, no scanner binaries. Plus this documentation and a least-privilege fix to the publish workflow.

0.3.0 2026-09-26

The cloudg rebrand

The project, package, CLI, Docker image, installers and CI became cloudg. Multi-cloud authentication (OIDC federation, STS assumption, managed identity, impersonation), compliance rulesets generated from Prowler's public data (28 frameworks, 4,166 controls, 10,236 check mappings), uv packaging, PyPI trusted publishing, GHCR image, MIT license. Patches folded in: 0.3.1 stopped the IaC scanners from silently scanning the working directory; 0.3.2 rekeyed dedupe on scanner check IDs with the cross-scanner equivalence map and moved the CLI to a Rich UI layer.