Cyber Knowledge · Domain 10 of 11 · Practitioner field guide

OSINT & Reconnaissance

OSINT is disciplined inquiry using lawfully accessible information. Reconnaissance is the structured discovery of entities, infrastructure, relationships, exposure, and change. This guide moves from questions and authority to reproducible collection, careful attribution, protected evidence, explicit confidence, and operational handoff—not from a name to an invasive search.

Authorization, privacy, and safety boundary

Publicly visible does not mean ethically unrestricted, risk-free, accurate, or authorized for active testing. Follow applicable law, platform terms, organizational policy, data-protection requirements, contractual scope, and a documented collection plan. Do not harass, impersonate, pretext, bypass access controls, purchase illicit data, expose vulnerable people, or actively probe systems without explicit authority. Use qualified legal and privacy review where the purpose, jurisdiction, or handling is uncertain.

Version 1.0 Published Source review: Status: maintained practitioner guide Maintained by Andrey Pautov Editorial policy and corrections

Orientation

Find less, understand more, and preserve the path between source and conclusion

OSINT quality is not measured by tabs opened, records collected, tools installed, or nodes drawn. It is measured by whether the collection answered a defined question, stayed inside authority and minimization boundaries, preserved provenance, separated observation from inference, survived challenge, and produced a proportionate decision.

Discovery

Identify candidate entities, infrastructure, aliases, services, publications, relationships, and changes from sources suited to the question. Record what each source actually observes and what it cannot establish.

Verification

Corroborate important claims through independent evidence, temporal consistency, source proximity, content analysis, and explicit alternatives. A repeated claim is not independent confirmation when every copy derives from one origin.

Protection

Protect investigators, subjects, sources, credentials, queries, workstations, case data, and downstream users. Minimize personal and sensitive data, control access, and avoid actions that create harm or tip off a monitored actor.

Communication

Deliver a claim-evidence matrix, timeline, entity model, confidence assessment, limitations, alternative hypotheses, and recommended collection or response. Make uncertainty visible instead of hiding it behind a polished graph.

Open-source information

Lawfully accessible material that may be observed, requested, purchased through legitimate channels, or obtained under an applicable public-access process. Accessibility, legality, terms, and ethical use remain context-dependent.

OSINT

An analyzed product derived from open-source information to address a defined requirement. Search output becomes intelligence only after evaluation, corroboration, interpretation, and communication.

Passive reconnaissance

Collection from existing third-party or supplied observations without intentionally sending assessment traffic to the target. “Passive” does not remove privacy, terms, attribution, or operational-security risks.

Active reconnaissance

Direct interaction with target-controlled systems or people to discover information. It requires explicit scope, rate and safety controls, source attribution, stop conditions, and coordination.

Operating model

One traceable chain from requirement to supported claim

Use a durable investigation chain: requirement → authority and risk assessment → hypotheses → collection plan → source capture → normalization → entity resolution → corroboration → analysis → confidence → review → dissemination → retention or deletion. Every transition should be inspectable.

LayerQuestionMinimum recordFailure to avoid
RequirementWhich decision must this investigation support?Requestor, decision, precise questions, priority, deadline, audience, acceptable uncertaintyCollecting everything about a person or organization “just in case”
AuthorityWhat may be collected, from where, by which method, and for how long?Purpose, legal/policy basis, scope, prohibited actions, privacy and safety constraints, approverAssuming public visibility authorizes interaction, republication, or indefinite retention
CollectionWhich source can observe the required fact?Source, query, account/tool, timestamp, result, coverage, cost, terms, errors, capture hashTreating an aggregator as the original source or a missing result as proof of absence
IdentityDo records refer to the same entity?Stable identifiers, temporal overlap, shared attributes, conflicts, confidence, alternativesMerging by name, email pattern, IP, logo, or shared hosting alone
AnalysisWhat is observed, inferred, assessed, and still unknown?Claim-evidence matrix, hypotheses, confidence, source quality, contradictions, gapsGraph proximity, AI prose, or repeated reporting presented as proof
HandoffWhat action is proportionate and reviewable?Finding, scope, owner, evidence links, caveats, sensitivity, next step, retentionDoxxing, over-attribution, automated enforcement, or publication without harm review

Learning path

Fourteen-module OSINT and reconnaissance practitioner sequence

Complete the modules in order when building a new capability. Experienced researchers can use them as a review map, but the dependencies remain: collection without authority creates risk, data without provenance cannot be defended, and attribution without alternatives is storytelling.

Authority, ethics, privacy, safety, and operational security

Module 01

Before searching, establish purpose, authority, scope, collection classes, prohibited methods, dissemination, retention, and stop conditions. Distinguish research about organization-controlled infrastructure from research about individuals. The latter can expose vulnerable people, families, employees, sources, investigators, or unrelated tenants, even when fragments are publicly visible.

Collection risk assessment

Assess legal and policy basis, platform terms, expected personal/special-category data, minors or vulnerable persons, source sensitivity, jurisdiction, necessity, proportionality, retention, reidentification risk, potential notification, and harm if the case file leaks. Record why each data class is needed.

Investigator OPSEC

Use organization-managed research identities where permitted, isolated browser profiles or workstations, separate credentials, strong MFA, controlled downloads, DNS and browser telemetry awareness, approved VPN/proxy design, malware-safe handling, logging, and a documented rule for when not to authenticate, click, download, contact, or join a community.

  • Passive is not invisible: search providers, websites, social platforms, archives, and APIs can log queries, source addresses, cookies, browser characteristics, and account identity.
  • Public is not ownerless: copyright, database rights, contractual terms, confidentiality, privacy, safeguarding, and contextual integrity may still apply.
  • Contact changes the operation: messaging a subject, requesting access, creating a pretext, or joining a restricted group can create legal, ethical, evidentiary, and safety consequences.
  • Data brokers are not neutral: provenance, consent, accuracy, access legitimacy, and downstream risk must be reviewed before acquisition or use.
Workflow: define purpose → classify target and affected people → determine authority → select minimally intrusive methods → approve research identity and environment → set prohibited actions and stop conditions → record dissemination and retention → brief the investigator → review when scope changes.
Evidence: collection authority, privacy/safety assessment, scope and exclusions, approved identities and tooling, data classes, contact prohibition, escalation path, retention schedule, and reviewer.
Boundary: this guide does not authorize investigation, access, surveillance, impersonation, social engineering, active scanning, or publication. Obtain context-specific approval. Do not use OSINT to facilitate stalking, harassment, credential abuse, physical targeting, or discrimination.
Accept when: an independent reviewer can tell what may be collected, why it is necessary, what must not be done, which harms are controlled, and when work must stop.

Primary: OHCHR and UC Berkeley — Berkeley Protocol on Digital Open Source Investigations · NIST SP 800-115

Intelligence requirements, hypotheses, and collection planning

Module 02

Translate a broad request into answerable questions. “Research this company,” “find everything about this person,” and “map the attack surface” have no natural stopping point. A defensible requirement names the decision, target entity, time window, relevant relationship, acceptable evidence, priority, constraints, and the uncertainty the decision-maker can tolerate.

Question decomposition

Break the requirement into claims: legal identity; owned domains and netblocks; externally operated services; brands and products; historical changes; supplier dependencies; public code; known malicious relationships; or whether one observable is connected to a campaign. State what would support, weaken, or falsify each hypothesis.

Collection matrix

For every question, list preferred original sources, secondary sources, query variants, identifiers, time coverage, expected result type, access and rate limits, privacy class, collection owner, validation route, fallback, and stop rule. Assign source priority before results bias the method.

Build the seed set without contaminating the conclusion

Start from known-good identifiers: official legal name, verified domain, registry identifier, supplied IP or hash, published account, or certificate fingerprint. Preserve where each seed came from. Expand one relationship at a time and label every edge by relationship type, observation time, and source. Never let a candidate discovered during collection silently become a confirmed seed.

Workflow: record the decision → write priority intelligence requirements → define scope and target identifiers → create competing hypotheses → choose observable indicators for each → map sources and query order → set collection limits → peer-review the plan → execute and update gaps without rewriting the original question.
Evidence: requirement record, target dictionary, hypothesis register, collection matrix, source register, query log, coverage and gap log, changes to scope, and completion criteria.
Boundary: a tool’s available fields should not define the investigation. Start with the question; use tools only where their observation model can contribute evidence.
Accept when: every planned query supports a named question, has a justified source and collection boundary, and can be stopped when sufficient evidence or a defined limit is reached.

1200km: Cyber Threat Intelligence field guide · AdversaryGraph IOC investigation workflow

Search engines, web archives, public records, and documents

Module 03

General search is a discovery layer, not a source of truth. Query multiple indexes because coverage, ranking, geography, language, personalization, removals, caching, and date interpretation differ. Capture the result page, query, locale, time, and destination. Open the underlying document and locate its publisher, edition, date, context, and revision history.

  • Query design: combine exact phrases, aliases, identifiers, site or domain restrictions, file types, date ranges, language variants, transliterations, known document titles, and exclusions. Keep a query notebook so another researcher can reproduce the path.
  • Archives: compare captures over time, but distinguish capture time from publication time. Missing snapshots do not prove a page never existed. Archived scripts, redirects, images, robots exclusions, and dynamic content may be incomplete.
  • Documents: inspect title, author, publisher, revision, embedded links, signatures, metadata, page references, and visual consistency. Sanitize downloads and avoid opening untrusted active content on normal workstations.
  • Public records: use the authoritative registry, court, procurement, corporate, regulatory, standards, or government portal where available. Record jurisdiction, identifier, effective date, status, and whether the database is complete or delayed.
  • Languages: preserve the original text and source. Treat machine translation as an aid, not a substitute for qualified review when nuance changes the conclusion.
Workflow: build alias and identifier variants → search multiple indexes → capture result and destination → pivot to original publisher or authoritative registry → retrieve safely → hash and store → extract entities and dates → compare archived versions → corroborate material claims → document absence limits.
Evidence: exact query, engine/source, timestamp, locale, result URL, final URL, HTTP/archival context where available, document hash, publisher, publication/revision date, relevant excerpt location, and capture method.
Boundary: cached snippets, AI summaries, search-result counts, autogenerated profile pages, and scraped mirrors are discovery aids. Do not cite them as original evidence when the authoritative item exists.
Accept when: each material claim links to the closest available original source, the exact supporting passage or record is preserved, and archive/search limitations are explicit.

1200km: Essential CLI tools for reconnaissance · Network reconnaissance guides

Domain, DNS, RDAP, certificate transparency, and routing research

Module 04

Infrastructure research works best as a time-aware evidence graph. A domain can resolve to shared CDN addresses; a certificate can cover customer-selected names; an ASN can announce provider and tenant space; and registration records can be privacy-protected or redacted. Every relationship needs a type, observation time, source, ownership interpretation, and confidence.

Names and registration

Use DNS for current records and authorized passive-DNS providers for historical observations. Use RDAP for structured domain, IP network, and autonomous-system registration data where supported. Preserve registrar/registry status, events, nameservers, entities returned, notices, redaction, and query time. Registration, administration, hosting, and operational control are different relationships.

Certificates and routing

Certificate Transparency logs can reveal names submitted in certificates and support monitoring for unexpected issuance. Inspect subject alternative names, issuer, validity, fingerprint, key and signature properties, source log, and first/last observation. Use RIR/RDAP and BGP sources for prefixes, origin AS, route history, and organization records while distinguishing allocation, announcement, hosting, and service ownership.

Resolution workflow

Start with the verified domain. Enumerate current authoritative DNS records and nameserver delegation. Search certificate records for exact and wildcard names, then validate candidates against current and historical DNS. Resolve IPs to registration and routing context. Identify CDN, cloud, shared-hosting, WAF, reverse proxy, email, SaaS, and third-party boundaries. A candidate becomes organization-attributed only when two or more suitable signals support that relationship or an authoritative source confirms it.

Workflow: canonicalize the name → query DNS and DNSSEC context → retrieve RDAP → inspect CT and live TLS → resolve and classify addresses → map RIR/ASN/BGP context → compare historical observations → identify shared-service boundaries → record candidates separately from confirmed assets.
Evidence: normalized FQDN, record type/value/TTL, first and last observation, resolver/provider, RDAP object and notices, certificate fingerprint and names, IP prefix, ASN, route observation, provider classification, and ownership rationale.
Boundary: a shared IP, wildcard certificate, reverse-DNS suffix, analytics identifier, or common nameserver does not by itself prove common ownership. CDN and SaaS observations often describe delivery infrastructure rather than the customer’s origin.
Accept when: confirmed assets are separated from candidates and third-party services, historical observations include time bounds, and every ownership statement identifies the evidence and alternative explanation.

Primary: ICANN RDAP · RFC 9162 — Certificate Transparency Version 2.0 · 1200km: theHarvester guide

Module 05

Internet-wide search platforms expose observations collected at particular times through particular scan methods. Shodan’s fundamental record is a service banner; Censys models hosts, web properties, certificates, and related assets; urlscan stores website-scan observations and resources. Their outputs are valuable evidence leads, but they are not live ground truth, vulnerability proof, ownership proof, or permission to test.

Query discipline

Begin with confirmed organization identifiers: netblocks, exact domains, certificate names or fingerprints, approved organization fields, or known products. Use structured filters, narrow time windows, and exported raw records. Record account tier and coverage because history, fields, results, quotas, and API features vary.

Observation interpretation

Separate IP host, virtual host, web property, service, certificate, and screenshot observations. Note observed-at time, scanner vantage, protocol, port, transport, banner, HTTP/TLS metadata, labels, inferred software, CPE/CVE claims, and provider. Validate critical exposure against an authorized current source before escalation.

  • Shodan: search banners and structured properties. Preserve the full banner and observation date; product or vulnerability fields remain scanner-derived claims.
  • Censys: choose hosts for IP and non-HTTP service questions, web properties for hostname/HTTP behavior, and certificates for issuance and presentation relationships.
  • urlscan: search historical scans before submitting a URL. Submission visibility matters; a public scan can disclose an investigation target or sensitive URL.
  • Reputation platforms: distinguish crowdsourced labels, vendor verdicts, passive associations, detections, and analyst comments. Count sources, not repeated feeds that inherit one label.
Workflow: validate seed → query exact records → export raw observations → normalize assets/services/timestamps → classify provider and shared tenancy → compare independent platforms → verify material exposure through authorized current inventory or scanning → open remediation with evidence and expiry.
Evidence: platform, dataset, query, account tier, timestamp, returned observation time, raw record, asset type, IP/hostname/port, service metadata, certificate, screenshot provenance, coverage caveat, and verification status.
Boundary: never equate an internet-index result with current exploitable exposure. A banner may be stale, deceptive, proxied, shared, incorrectly parsed, or changed since observation. Do not trigger an on-demand scan without authority.
Accept when: every exposure is tied to a time and dataset, shared infrastructure is handled correctly, and high-impact findings are revalidated through an authorized current source.

Official: Shodan query fundamentals · Censys Platform quick start · urlscan Search API · 1200km: Shodan guide · Censys guide

Web, API, mobile, and application reconnaissance

Module 06

Application reconnaissance models entry points, trust boundaries, identities, data flows, technologies, deployments, and third parties. Passive discovery can include indexed pages, archived content, public documentation, application stores, package registries, public API specifications, DNS, certificates, client-side code already delivered to a normal user, and exposure-index observations. Direct crawling, directory enumeration, fingerprinting, parameter discovery, or protocol interaction is active testing and must stay inside explicit scope.

  • Application inventory: hostname, scheme, ports, environment, product, owner, authentication, user roles, API base URLs, mobile package identifiers, status pages, documentation, repositories, data classes, and third-party integrations.
  • Client-visible surface: links, scripts, source maps, API endpoints, feature flags, telemetry endpoints, storage names, CSP/connect destinations, public configuration, and framework metadata. Do not treat an endpoint name as proof that it is reachable or vulnerable.
  • API context: published OpenAPI/GraphQL documentation, versioning, base paths, authentication scheme, rate-limit statements, SDKs, examples, deprecation, status, and public change logs.
  • Historical surface: archived routes, old documentation, retired subdomains, mobile versions, abandoned marketing applications, and previous hosting. Validate current ownership and availability before classifying.
  • Active phase: use an approved user agent, source address, rate, time window, request method set, wordlist, exclusion list, authentication account, data-handling rule, and emergency contact.
Workflow: establish owned application seeds → enumerate passive sources → build endpoint and trust-boundary model → reconcile inventory and owners → mark candidates → approve active phase if needed → collect low-impact confirmations → stop on instability or out-of-scope redirection → report exposure separately from vulnerability.
Evidence: source URL, capture time, application/environment, endpoint, method if known, authentication context, observed client reference or response, archive status, scope state, owner, and validation status.
Boundary: do not fuzz, brute-force, enumerate accounts, bypass controls, upload payloads, or probe third-party integrations under the label “OSINT.” Those are active security tests and require their own authority and safeguards.
Accept when: the application surface distinguishes observed, inferred, historical, third-party, and actively verified components, with owners and scope status attached.

Primary: OWASP Web Security Testing Guide · 1200km: Web application reconnaissance · Secure Code / Application Security

People, organizations, brands, suppliers, and relationship research

Module 07

People and organization research carries the highest risk of unnecessary collection and harm. Define the organizational question first: confirm an executive role from authoritative publications, identify the owner of a business service, validate a supplier relationship, detect brand impersonation, or understand public hiring signals relevant to a threat model. Avoid building personal dossiers where an organization-level answer is sufficient.

Authoritative organization sources

Prefer official corporate sites, filings, registries, regulator notices, procurement records, verified press releases, conference programs, patent or standards records, repositories, and the subject’s own professional publication. Record effective dates; titles, subsidiaries, brands, and supplier relationships change.

Identity resolution

Use multiple stable attributes: legal entity or employer, verified account, organization domain, location at the relevant time, publication history, unique identifier, or explicit cross-link. Names, profile photos, usernames, writing style, and contact patterns are weak alone. Preserve conflicts and avoid inferring sensitive attributes.

  • Collect the minimum professional information necessary for the requirement and mask personal contact details in routine reporting.
  • Distinguish current employee, former employee, contractor, speaker, contributor, customer, partner, reseller, and unverified self-claim.
  • For brand monitoring, preserve the legitimate brand baseline—domains, certificates, visual marks, support channels, and campaigns—before labeling a lookalike.
  • Do not contact colleagues, relatives, or unrelated third parties; do not attempt password recovery, account discovery, or identity challenge flows.
  • Escalate safeguarding concerns and credible threats through approved channels rather than publishing personal information.
Workflow: state organization-level question → minimize attributes → collect official records → add self-published professional sources → resolve identity with independent signals → record temporal validity → review privacy and harm → redact outputs → retain only what supports the decision.
Evidence: entity type, exact name, stable identifier, role/relationship, source publisher, direct URL, publication/effective date, capture time, corroboration, conflicts, sensitivity, redaction, and confidence.
Boundary: never infer ethnicity, religion, health, sexuality, political belief, home location, family relationship, or other sensitive characteristics unless a lawful and necessary investigation explicitly requires qualified handling. Public availability does not make inference accurate or proportionate.
Accept when: the answer uses the least personal data possible, material identities and relationships are time-bounded and corroborated, and reporting cannot reasonably expose unrelated people.

Primary: Berkeley Protocol · 1200km: Privacy / Data Handling · Privacy and data governance

Public code, packages, cloud artifacts, documents, and exposed secrets

Module 08

Public repositories and package ecosystems can explain technologies, ownership, deployment patterns, identifiers, domains, storage names, API schemas, historical changes, and third-party dependencies. They can also expose credentials or sensitive configuration. Treat suspected secrets as hazardous evidence: do not test them, do not place them in tickets or chat, and do not broaden access.

  • Repository discovery: search verified organizations, repository topics, package metadata, commit authorship where appropriate, documentation links, dependency files, CI workflows, infrastructure templates, container references, public issue trackers, and release artifacts.
  • Attribution: a repository name, copied code, contributor email, package namespace, or cloud identifier may be stale, forked, mirrored, typosquatted, or unrelated. Confirm through organization ownership, signed releases, official links, registry publisher data, and temporal consistency.
  • Secret handling: preserve only the minimum needed to report; mask values; record repository, path, commit, discovery time, type, exposure status, and responsible contact. Do not clone unnecessary history if a safer permalink or platform report is sufficient.
  • Cloud artifacts: bucket names, tenant IDs, account IDs, regions, image names, workload identity, and endpoints are identifiers, not proof of public access. Validate ownership and exposure through the owner or an explicitly authorized assessment.
  • Dependencies: package name and version may support technology inventory, but lockfiles, built artifacts, runtime inventory, and deployment evidence provide different confidence levels.
Workflow: identify verified organization and packages → query narrow identifiers → capture permalink/commit metadata → classify source and temporal validity → route suspected secrets privately without use → reconcile software/cloud clues with authoritative inventory → document candidate relationships → delete unnecessary sensitive copies.
Evidence: platform, organization/repository/package, immutable commit or release identifier, path, timestamp, publisher, signature/verification state, relevant masked excerpt, secret-handling record, ownership status, and remediation handoff.
Boundary: discovering a credential does not authorize authentication. Do not verify by logging in, querying an API, listing a bucket, or attempting a transaction. Follow the platform’s reporting mechanism and the organization’s incident process.
Accept when: repository and package attribution is supported, sensitive values are masked and access-controlled, no credential was used, and every inferred asset remains a candidate until owner evidence confirms it.

Official: GitHub repository visibility and security considerations · 1200km: Cloud Security · Secure Code / Application Security

Images, video, audio, geolocation, chronolocation, and media verification

Module 09

Media verification asks whether an item is original or derivative, what it depicts, where and when it may have been captured, how it was published, and whether manipulation or context changes the claim. No single metadata field, reverse-image result, shadow estimate, AI detector, compression artifact, or visual resemblance should carry the conclusion.

Preserve before analysis

Capture the original URL, publisher/account, post and retrieval times, surrounding text, engagement context, available original file, hashes, file metadata, dimensions, codec/container, thumbnails, and archive reference. Work from copies and record transformations such as transcoding, frame extraction, cropping, enhancement, or translation.

Corroborate location and time

Compare landmarks, terrain, road geometry, signs, scripts, utilities, transit, vegetation, weather, sun direction, shadows, construction state, event schedules, satellite or map imagery, and independent media. Document which features are distinctive, which are common, and which alternatives remain possible.

  • Use reverse-image and keyframe searches across multiple indexes to locate earlier publication and variants.
  • Separate file creation metadata from upload, edit, transcoding, screenshot, and publication times.
  • Do not expose precise locations of vulnerable people, shelters, medical facilities, witnesses, or ongoing operations.
  • Track source lineage: original uploader, reposts, crops, translations, overlays, compilations, and news embeds can create false independence.
  • Treat automated deepfake or manipulation scores as one signal with documented model/version/limitations, not a verdict.
Workflow: preserve item and context → hash originals → map publication lineage → extract frames/metadata safely → search earlier variants → form location/time hypotheses → test distinctive features against independent sources → document contradictions → peer review → publish only the precision necessary.
Evidence: original and capture URLs, account/publisher, timestamps, hashes, tool/version, transformations, keyframes, feature table, map/weather/archive references, alternative locations/times, confidence, and safeguarding decision.
Boundary: avoid claiming exact location or time when only country, region, season, or broad period is supported. Never sharpen uncertainty into false precision because a map pin or timestamp field is convenient.
Accept when: the conclusion can be reconstructed from preserved media and independent features, source lineage is understood, alternatives were tested, and publication precision passes a harm review.

Primary: Berkeley Protocol on Digital Open Source Investigations

Threat infrastructure, IOC enrichment, campaign pivots, and attribution restraint

Module 10

Threat-infrastructure research begins with a supplied observable or behavior and asks which relationships are supported: resolution, hosting, certificate presentation, content similarity, registration, malware communication, file download, redirect, tracker reuse, or co-observation in a report. Shared infrastructure and commodity services make careless clustering especially dangerous.

Observable discipline

Canonicalize IP, domain, URL, email, file hash, certificate fingerprint, mutex, registry path, user agent, or other artifact. Record source, first/last seen, role, confidence, TLP/handling, maliciousness basis, and expiration. Preserve exact URLs safely without browsing a live malicious destination from a normal workstation.

Relationship discipline

Use explicit predicates: domain resolved-to IP at time; certificate presented-by host; file contacted domain in sandbox run; URL redirected-to URL; report asserted actor-used infrastructure; registration shared email; content hash matched. “Related to” is too vague for analytic review.

Attribution ladder

Move from observation to cluster only when multiple independent or high-quality relationships align across time, infrastructure, behavior, malware, operations, and victimology. Keep infrastructure clustering, activity-cluster naming, malware-family assessment, campaign assessment, and sponsor attribution as separate conclusions. A TTP overlap, shared hosting provider, common registrar, commodity tool, or AI-generated similarity score is a lead—not attribution.

Workflow: validate and normalize seed → query local and approved external sources → capture direct relationships and time bounds → distinguish malicious, suspicious, benign, sinkhole, scanner, shared, and unknown roles → cluster with explicit rules → compare behavior and reporting → document alternatives → expire unsupported IOCs → hand off supported TTP and detection leads.
Evidence: observable type/value, source, handling, first/last seen, direct relationships, source independence, provider and shared-tenancy context, malware/report evidence, ATT&CK mapping rationale, confidence, expiry, and contradictions.
Boundary: reputation is context, not identity. An IP can be reassigned; a domain can be sinkholed; a certificate can be shared; and a public scanner can appear in victim telemetry. Confirm time and role before blocking or attribution.
Accept when: each graph edge has a typed, dated source; cluster rules are reproducible; generic signals are discounted; and attribution strength does not exceed the evidence.

Official: VirusTotal relationship model · MITRE ATT&CK Reconnaissance tactic · 1200km: From log to report with AdversaryGraph · AdversaryGraph IOC investigation capabilities

Automation, APIs, entity graphs, normalization, and data engineering

Module 11

Automation should make collection reproducible and reviewable, not turn every weak association into a permanent fact. Design pipelines around immutable raw capture, normalized records, typed relationships, provenance, temporal validity, confidence, access control, review state, and deletion. Preserve the ability to rebuild derived graphs when parsing or resolution logic changes.

  • Raw layer: source response, request metadata, collection timestamp, tool and version, account tier, status/error, response headers where relevant, content hash, encryption, and retention class.
  • Normalized layer: canonical entity type and value, stable internal ID, source ID, observation time, valid-from/to, labels, sensitivity, and parser version. Keep raw values for audit.
  • Relationship layer: subject, predicate, object, source, observation time, confidence, inference rule, reviewer, status, and expiry. Avoid generic edges.
  • Analytic layer: hypotheses, claim-evidence links, source independence, scoring explanation, contradictions, gaps, and review history.
  • Operational layer: case, owner, dissemination, alert/detection/remediation links, suppression, retention, and deletion proof.

Reliability engineering

Honor rate limits, quotas, acceptable-use terms, pagination, time zones, retries, backoff, schema changes, partial responses, and API deprecation. Cache according to source terms and freshness needs. Deduplicate by source-native ID plus canonical observation—not by value alone—so multiple independent sources and historical changes remain visible. Quarantine malformed or ambiguous records instead of silently coercing them.

Workflow: register source and terms → define schema and mappings → collect into immutable raw storage → validate and normalize → resolve candidates without destructive merge → create typed edges → run confidence and expiry rules → queue analyst review → publish approved facts → monitor data quality and rebuild when logic changes.
Evidence: source registry, API/version, schema, raw hash, parser version, validation result, normalization rule, dedup key, entity-resolution decision, relationship rule, analyst approval, quality metrics, and retention action.
Boundary: never hide source errors, missing pages, quota truncation, or parser failures behind a successful job status. A partial dataset cannot support an unqualified negative finding.
Accept when: a derived entity and every relationship can be traced to raw observations, uncertain merges remain reversible, failed records are visible, and deletion propagates through derived products.

Official tooling: OWASP Amass · urlscan search semantics · 1200km: AdversaryGraph platform guide

AI-assisted OSINT with bounded tools, citations, and human review

Module 12

Apply the AI Security agent and MCP control model before connecting a model to search, enrichment, browser, scanning, or publication tools.

AI can accelerate query generation, translation, document triage, entity extraction, schema normalization, relationship suggestions, timeline drafting, hypothesis generation, and report formatting. It can also fabricate sources, merge identities, expose sensitive case data, follow malicious content instructions, overstate confidence, and initiate actions outside scope when connected to tools.

Safe task design

Provide a structured requirement, approved source set, data classification, target schema, prohibited inferences, citation requirement, confidence vocabulary, and instruction to return unknown when evidence is absent. Use retrieval over approved case material and require source spans for every proposed claim.

Tool and MCP controls

Expose read-only, allowlisted tools by default. Separate search from active scanning, authentication, messaging, purchase, upload, and publication. Require explicit human authorization for any state-changing or target-interacting action. Log tool name, arguments, source, response hash, model, prompt version, and reviewer decision.

  • Redact unnecessary personal data, secrets, private sources, and customer information before remote processing; follow provider data-handling policy.
  • Treat webpages, documents, repository content, images, and search results as untrusted data that may contain prompt-injection instructions.
  • Resolve entities deterministically where possible. AI may suggest a merge, but a reviewer must inspect supporting and conflicting identifiers.
  • Recalculate dates, IP ranges, hashes, DNS names, geographic coordinates, and query syntax with deterministic validation.
  • Preserve citations to the original item, not to the model response or a generated summary.
  • Evaluate on known cases for source fidelity, identity precision/recall, abstention, privacy leakage, unsafe tool selection, and confidence calibration.
Workflow: classify case data → choose permitted provider/local model → retrieve only approved evidence → run a versioned task-specific prompt → validate schema → resolve citations → verify entities and calculations → challenge with alternatives → review every proposed claim/action → save model and reviewer audit trail.
Evidence: provider/model, prompt and policy version, retrieval set, tool calls, output schema validation, citations, redactions, reviewer edits/rejections, confidence, latency/cost where relevant, and evaluation result.
Boundary: never let an agent independently expand a person-focused target, contact a subject, authenticate, submit a URL for public scanning, buy data, trigger active reconnaissance, or publish an assessment. AI output is a lead until the underlying evidence is reviewed.
Accept when: the system abstains appropriately, every accepted claim resolves to reviewed evidence, tool authority is enforceable, sensitive data follows policy, and a human owns the final assessment.

1200km: AI & Offensive Security research · HexStrike laboratory index · Shodan with HexStrike AI · AI-assisted OSINT exposure-map case study

Defensive external attack-surface discovery and continuous monitoring

Module 13

A defensive exposure program reconciles what the organization believes it owns with what external observers can see. Its goal is not a one-time internet scan. It maintains an attributed inventory of domains, certificates, IPs, netblocks, cloud and SaaS properties, applications, APIs, repositories, mobile apps, brands, and suppliers; detects meaningful change; verifies ownership; routes remediation; and measures closure.

Discovery-to-inventory loop

Seed from legal entities, brands, known domains, cloud accounts, certificates, netblocks, repositories, and service owners. Discover candidates through approved sources. Score ownership evidence separately from exposure risk. Ask the responsible owner to confirm or reject. Add confirmed assets with lifecycle, criticality, environment, dependency, and monitoring coverage.

Change-to-action loop

Track first/last seen, new certificate, DNS change, new port/service, login or administration page, exposed development environment, new repository/package, storage reference, security-header change, or brand lookalike. Deduplicate, suppress expected provider churn, verify current state, and route by owner and severity.

Metrics that support decisions

Measure candidate-to-confirmed precision, unknown-owner backlog, time to ownership decision, time to verification, exposure age, stale assets, orphaned domains, unexpected certificates, high-risk changes, remediation age, false-positive causes, source freshness, and coverage by critical service. Avoid using the raw number of discovered assets or scanner “risk score” as a maturity grade.

Workflow: synchronize authoritative inventory → collect external observations → normalize and cluster → classify shared/third-party infrastructure → score ownership evidence → obtain owner confirmation → validate exposure safely → create finding and response → rescan/requery for closure → update inventory and detection logic.
Evidence: authoritative asset ID, candidate identifiers, ownership signals, provider boundary, first/last seen, current verification, service/technology, finding, owner, remediation, closure proof, source coverage, and monitoring rule.
Boundary: third-party SaaS, CDN, cloud-provider, shared-hosting, and acquired-company assets require explicit relationship modeling. Do not import every hostname on a shared IP or certificate into company inventory.
Accept when: new external observations reliably reach an accountable owner, candidates are not silently promoted, current verification is proportionate, and closure updates both inventory and monitoring.

Primary: CISA BOD 23-01 (mandatory only for its stated FCEB scope; useful visibility concepts elsewhere) · 1200km: AdversaryGraph Asset Surface guide · Cloud posture and exposure

Analysis, confidence, evidence preservation, reporting, and operational handoff

Module 14

The final product should make the reasoning inspectable. Separate facts observed directly, source assertions, analytic inferences, and unknowns. Use consistent confidence language tied to evidence quality, source independence, identity resolution, temporal relevance, contradictions, and alternative explanations. Confidence is not the same as impact, priority, or probability.

Claim-evidence matrix

For each claim, record supporting observations, contradicting observations, source proximity, independence, timestamps, identity rationale, assumptions, alternative hypotheses, confidence, and reviewer. Link graphs, timelines, and screenshots back to the matrix rather than letting visuals become unsupported conclusions.

Preservation package

Maintain original capture, normalized record, query and collection log, hashes, timestamps, tool/version, transformations, access and handling, analyst notes, source restrictions, review history, and export. For legal or disciplinary use, obtain qualified evidence-handling guidance and preserve chain-of-custody requirements.

  • Executive output: question, supported answer, confidence, material implications, limitations, decision options, and next review.
  • Operational output: confirmed entities/observables, scope, owner, timeline, typed relationships, source links, detection/remediation actions, expiry, and handling.
  • Analytic annex: methodology, sources, queries, identity resolution, alternatives, gaps, contradictions, AI/tool use, and quality controls.
  • Redaction: remove personal contact data, credentials, irrelevant private details, precise safeguarding-sensitive locations, and provider secrets.
  • Correction: provide a route to challenge or correct a record; preserve what changed, why, by whom, and which downstream products were updated.
Workflow: freeze collection set → normalize final entities → build timeline and typed graph → complete claim-evidence matrix → test alternatives → assign confidence → peer review identity/privacy/method → produce audience-specific report → record dissemination → track action/correction/expiry → delete data when retention ends.
Evidence: complete case manifest, original and derived hashes, query log, source register, claim-evidence matrix, entity-resolution decisions, confidence rubric, alternative hypotheses, redaction log, reviewer approval, dissemination, and retention event.
Boundary: write “not observed in the searched sources and time window,” not “does not exist.” Write “assessed with moderate confidence,” not “confirmed,” when the relationship depends on inference.
Accept when: another analyst can reproduce the collection path, distinguish evidence from inference, evaluate alternatives, understand limitations, and execute the handoff without receiving unnecessary sensitive data.

1200km: External validation and evidence · DFIR evidence and timeline practice · AdversaryGraph campaign-cluster investigation

Source strategy

Choose sources by observation model, not brand familiarity

No platform “has all the data.” Select a source because it can observe the entity or relationship needed, then document coverage, timing, terms, and validation. The matrix below is a decision aid, not an endorsement or claim of completeness.

QuestionPreferred source classWhat it can supportEssential caveat
Who is the registered domain/IP/ASN entity?Registry/RIR RDAP and authoritative registration servicesStructured registration object, status, events, entities returned, noticesRedaction and privacy proxies are common; registration is not operational ownership
What does a name resolve to now or historically?Authoritative/current DNS and approved passive-DNS providersDated record observation and resolution historyResolver cache, split horizon, sinkholes, wildcarding, and provider coverage affect results
Which names appeared on certificates?CT logs, certificate datasets, live TLS where authorizedCertificate issuance/logging and presentation observationsA certificate name or wildcard does not prove live service or organization ownership
Which public services were observed?Internet-wide host/web-property search platformsDated banner, service, protocol, web, TLS, or screenshot observationsRecords may be stale, shared, parsed incorrectly, or different from current state
What pages/resources were loaded?Website-scan archives and owner-authorized browsing/captureDated page, request, response, resource, redirect, and visual contextSubmitting a new scan may interact with or publicly disclose the target
What did the organization publish?Official site, filings, registry, release, documentation, repositoryFirst-party assertion at a time and versionSelf-published claims still require interpretation and may later change
How is an observable related to malware?Sandbox/report evidence and typed threat-intelligence relationshipsObserved communication, download, execution, or asserted report relationshipReputation aggregation and shared infrastructure can produce misleading associations
What existed in the past?Web archives, historical DNS/cert/service records, version controlPoint-in-time capture or change historyArchive coverage is incomplete; absence of capture is not absence of existence
Does an asset currently belong to us?Authoritative inventory, account/provider data, owner confirmationCurrent accountable ownership and lifecycleExternal observation alone rarely establishes full ownership

Reusable artifacts

Templates that preserve method, evidence, and restraint

A tool output is not an investigation record. These artifacts connect collection to decision and make the work reviewable, correctable, and repeatable.

Investigation charter

Case ID; requestor; decision; priority questions; target entity and known-good seeds; purpose and authority; jurisdictions; approved source/method classes; prohibited actions; personal/sensitive data; safeguarding; research identity; active-contact rule; dissemination; retention; owner; approver; start/stop conditions.

Collection record

Collection ID; question; source and dataset; exact query/request; account/tool/version; collection and observation time; locale/time zone; result and pagination; raw hash; source terms/classification; errors/coverage; extracted entities; analyst; review status.

Entity record

Internal ID; entity type; canonical value/name; aliases; stable external identifiers; source observations; valid-from/to; ownership/role; sensitivity; candidate/confirmed/rejected state; merge and split rationale; conflicts; confidence; reviewer; expiry.

Typed relationship record

Subject; predicate; object; direct observation or inference; source; observed-at; valid time; first/last seen; supporting and contradicting evidence; shared-service context; inference rule; confidence; analyst/reviewer; expiry; downstream use.

Claim-evidence matrix

Claim; decision relevance; supporting evidence; contradictions; source proximity and independence; temporal fit; identity fit; alternative hypotheses; assumptions; gaps; confidence; wording; reviewer; correction status.

Exposure candidate

Candidate asset; discovery source; ownership signals; provider/tenant boundary; environment; service; current verification; risk lead; owner; confirmation status; first/last seen; remediation; closure evidence; inventory and monitoring update.

AI-assistance audit

Task; provider/model; prompt/policy; retrieval set; redaction; tool calls; output schema; citations; validation; unsafe or rejected suggestions; reviewer changes; accepted claims; cost/latency where material; evaluation version.

Case preservation manifest

Original files/URLs; capture method; hashes; timestamps; transformations; storage; access history; handling; source restrictions; derived files; reports; dissemination; legal hold where authorized; correction; retention trigger; deletion proof.

End-to-end practice

Six OSINT and reconnaissance case studies

These cases are designed for owned infrastructure, supplied datasets, documentation, or approved lab material. They demonstrate method and evidence handling without encouraging unscoped targeting of real people or third-party systems.

Case 1 — Reconcile a company’s external web estate

Purpose: defensive asset inventory · Inputs: legal entities, verified domains, CMDB export, approved external datasets

  1. Create known-good seeds from legal and inventory records; do not begin from an internet-wide fuzzy organization search.
  2. Collect DNS, RDAP, CT, approved passive DNS, host/web-property, repository, status-page, and application-store observations.
  3. Normalize FQDN, certificate, IP, ASN, repository, application, and provider entities; retain observation times.
  4. Classify direct ownership, acquired entity, SaaS, CDN/WAF, supplier, shared hosting, historical, candidate, or rejected.
  5. Ask owners to confirm candidates and environment; compare with cloud/account inventory and service maps.
  6. Verify material exposure through approved current sources; create owner-bound remediation and monitoring records.
Done when: every confirmed asset has an accountable owner and lifecycle, every candidate has an evidence-based disposition, and shared infrastructure has not been imported as company property.

Case 2 — Investigate a suspected phishing domain

Purpose: brand and threat triage · Inputs: supplied domain or URL, approved isolated collection environment

  1. Canonicalize the indicator and preserve the reporting source without opening it from a normal browser.
  2. Search existing URL-scan, DNS, RDAP, CT, reputation, certificate, redirect, screenshot, and historical observations before any new submission.
  3. Compare string, visual, certificate, infrastructure, content, and timing with the legitimate brand baseline.
  4. Identify shared hosting/CDN, registrar, nameserver, campaign-related infrastructure, and malware/report evidence without equating provider reuse with actor identity.
  5. Assign role and confidence; route domain/URL and supporting evidence to blocking, takedown, fraud, legal, or monitoring owners.
  6. Set expiry and recheck; record sinkhole, takedown, content, DNS, and certificate changes.
Done when: the maliciousness assessment is based on observed content/behavior or trusted reporting, brand similarity is explained, and the report does not over-attribute a provider or actor.

Case 3 — Validate an unexpected certificate alert

Purpose: certificate and domain monitoring · Inputs: CT alert, certificate record, organization domain inventory

  1. Preserve certificate fingerprint, SANs, issuer, validity, CT log/source, and alert time.
  2. Classify exact, wildcard, lookalike, internal-looking, customer-controlled, provider, or unrelated names.
  3. Query DNS and authorized historical sources; compare known certificate-management providers and account records.
  4. Contact the named certificate/domain owner through internal channels; do not interact with an unknown endpoint solely because it appeared in CT.
  5. Determine authorized issuance, abandoned DNS, provider validation artifact, brand abuse, or false candidate.
  6. Update inventory, issuance monitoring, CAA/control review, response, and closure evidence.
Done when: ownership and issuance status are confirmed by authoritative records or the accountable owner, and the CT observation remains distinct from live-service exposure.

Case 4 — Triage a public-repository secret report

Purpose: exposure response · Inputs: repository permalink or platform alert · Constraint: never use the credential

  1. Restrict access to the report; preserve repository, commit, path, time, secret type, and a masked excerpt.
  2. Confirm repository and organization relationship through authoritative links; identify fork, mirror, archived, or unrelated status.
  3. Notify the security/secret owner through approved private channels and revoke/rotate before broader investigation.
  4. Use platform history and internal provider logs—not authentication attempts—to determine validity and use.
  5. Remove or rewrite history where appropriate, enable detection and push protection, and inspect related exposure under incident authority.
  6. Delete unnecessary copies; document revocation, remediation, monitoring, and disclosure handling.
Done when: the value is revoked or proven non-sensitive through owner evidence, no investigator used it, exposure history is understood, and preventative controls are updated.

Case 5 — Build a threat-infrastructure cluster from one IOC

Purpose: CTI and detection support · Inputs: sourced IOC with time and handling · Output: typed graph, not actor attribution

  1. Validate type, source, role, first/last seen, handling, and why the seed is malicious or suspicious.
  2. Collect direct DNS, certificate, hosting, URL, redirect, sandbox, file, report, and registration relationships.
  3. Normalize time and separate shared, benign, scanner, sinkhole, CDN, and unknown infrastructure.
  4. Build candidate clusters using documented rules; require multiple high-quality signals for promotion.
  5. Compare behaviors and report evidence; map ATT&CK only from described behavior, not infrastructure type.
  6. Create detection/blocking/hunting leads with confidence and expiry; retain alternatives and contradictions.
Done when: every edge is typed and dated, cluster membership is reproducible, generic infrastructure is discounted, and no actor claim exceeds the reporting evidence.

Case 6 — AI-assisted review of a supplied research package

Purpose: analyst acceleration · Inputs: approved documents and captures · No autonomous collection or target interaction

  1. Classify and redact the package; create immutable originals and a retrieval corpus limited to approved evidence.
  2. Define schemas for entities, observations, typed relationships, claims, citations, contradictions, and unknowns.
  3. Ask the model to extract and propose—not decide—while treating source content as untrusted data.
  4. Validate identifiers, dates, citations, translations, calculations, and entity merges deterministically.
  5. Have an analyst review every claim and compare alternative hypotheses; reject uncited or overconfident output.
  6. Save provider/model/prompt/retrieval/tool/reviewer audit and generate the report from accepted structured records.
Done when: all accepted claims resolve to source spans, no tool exceeded authority, sensitive data followed policy, and the final assessment is owned by a named analyst.

Hands-on curriculum

Twelve-lab OSINT and reconnaissance sequence

Use your own domain, an intentionally published lab domain, reserved examples such as example.com, supplied offline datasets, and platform documentation. Do not redirect these exercises at an arbitrary real person or organization.

Lab 01 — Investigation charter

Turn “map our exposure” into three precise questions. Write scope, authority, source classes, prohibited methods, privacy risks, research identity, active-contact rule, retention, completion criteria, and escalation. Peer-review the charter before collecting.

Lab 02 — Query and provenance notebook

Use a harmless subject to run exact-phrase, site, file-type, date, language, alias, and identifier queries across two search engines. Preserve query, locale, time, result, original source, archive state, and what each missing result cannot prove.

Lab 03 — DNS/RDAP/CT evidence graph

For a domain you control, collect current DNS, ICANN/RIR RDAP, CT records, live certificate metadata, and provider context. Build typed dated edges and mark authoritative ownership, service provider, candidate, and historical relationships separately.

Lab 04 — Internet-index comparison

Compare approved Shodan, Censys, and urlscan observations for your own asset. Record dataset, query, observation time, host versus web-property semantics, stale/shared results, and current verification. Explain every disagreement.

Lab 05 — Passive application map

Map your own public website from normal browser-delivered HTML, scripts, CSP destinations, documentation, status pages, archives, DNS, and certificates. Separate observed endpoints from inferred or historical candidates; perform no fuzzing.

Lab 06 — Repository and package inventory

Use a repository you own. Trace organization, packages, releases, dependency files, container images, CI workflows, domains, and cloud identifiers. Create a private synthetic secret example and practice masked reporting without using the value.

Lab 07 — Media verification worksheet

Use an instructor-supplied non-sensitive image and known answer. Preserve file/context, extract metadata and frames, identify distinctive features, search variants, test two location/time hypotheses, and document confidence and transformation history.

Lab 08 — IOC relationship graph

Use a historical, defanged IOC package supplied with source records. Normalize entities; type DNS, certificate, URL, file, and report edges; identify shared infrastructure; create candidate clusters; add expiry; and write why the graph does not prove attribution.

Lab 09 — Source/API ingestion

Ingest a small saved JSON response. Keep raw hash and request metadata, validate schema, normalize observations, preserve source IDs, quarantine malformed rows, deduplicate without losing independent sources, and rebuild the graph from raw data.

Lab 10 — AI extraction evaluation

Provide an approved offline report with known entities and citations. Evaluate extraction precision/recall, identity merges, citation resolution, abstention, prompt-injection resistance, sensitive-data handling, and analyst correction. No external tools are enabled.

Lab 11 — External asset reconciliation

Combine a synthetic CMDB with DNS, certificate, and service observations containing shared-provider traps. Score ownership evidence, request simulated owner dispositions, confirm inventory changes, route exposures, and measure candidate precision and unknown-owner backlog.

Lab 12 — Defensible final product

Produce a charter, source/query register, entity dictionary, typed graph, timeline, claim-evidence matrix, confidence assessment, alternatives, redacted executive summary, operational handoff, correction route, and retention action. A second analyst must reproduce one claim.

Failure atlas

Common ways OSINT becomes unsafe or analytically weak

  • Collection without a decision“Find everything” creates uncontrolled scope, privacy risk, and a pile of facts no one can use. Rewrite the requirement and define stopping conditions.
  • Public-equals-permitted thinkingVisibility does not authorize active probing, contact, republication, or indefinite retention. Recheck authority, terms, necessity, and harm.
  • Name-only identity mergeCommon names and reused usernames create false people and organizations. Require stable corroborating identifiers and preserve alternatives.
  • Shared-infrastructure ownershipA CDN IP, wildcard certificate, cloud ASN, nameserver, or analytics ID pulls unrelated tenants into the graph. Model provider relationships explicitly.
  • Timestamp collapseCollection, publication, observation, registration, certificate validity, archive capture, and event time become one “date.” Preserve each semantic time separately.
  • Aggregator circularitySeveral platforms repeat one original report, creating apparent corroboration. Trace lineage and count independent origins.
  • Absence as proofNo search result becomes “does not exist.” State source set, window, coverage, errors, and “not observed.”
  • Stale exposure escalationAn old banner or screenshot becomes an urgent current vulnerability. Verify ownership and current state through an authorized source.
  • Credential verificationAn investigator tests a discovered secret. Revoke and investigate through authorized logs; never authenticate to “confirm.”
  • Public URL submissionA sensitive target is submitted to a public scanner, exposing the investigation. Search existing records first and select visibility deliberately.
  • Graph-as-proofLayout, color, centrality, or node proximity replaces typed evidence. Make every edge reviewable and keep candidate status visible.
  • AI citation launderingGenerated prose cites a search snippet or nonexistent source. Resolve every citation to original captured evidence before acceptance.
  • Unbounded agent toolsAn AI system expands scope, scans, contacts, or publishes. Enforce allowlists, read-only defaults, approval gates, logging, and stop controls.
  • Investigator exposureNormal accounts, browser profiles, or workstations reveal identity or ingest malicious content. Use approved research identities and isolated handling.
  • Overcollection of peoplePersonal details enter a company-security case without necessity. Minimize, redact, restrict, and delete; escalate safeguarding issues.
  • No correction pathA false association persists across graphs, reports, alerts, and blocklists. Version claims and propagate correction or deletion downstream.

Readiness gate

Accept an OSINT product only when the method and limits are visible

  • The product names the decision, requestor, questions, scope, time window, and completion criteria.
  • Authority, privacy, platform terms, investigator safety, source handling, contact rules, and retention were reviewed.
  • Known-good seeds and candidate discoveries remain distinguishable throughout the case.
  • Every material observation records source, query or retrieval route, collection time, observation/publication time where available, and coverage caveat.
  • Original captures are hashed or otherwise integrity-protected; transformations and tools are documented.
  • Entities were resolved using appropriate stable identifiers; weak name, IP, certificate, or hosting matches were not silently merged.
  • Relationships use explicit predicates, evidence, temporal bounds, confidence, and expiry rather than generic “related” edges.
  • Original, secondary, and aggregate sources are distinguished; apparent corroboration was checked for common lineage.
  • Facts, source assertions, inferences, assumptions, and unknowns are visibly separated.
  • Alternative hypotheses and contradicting evidence were recorded and reviewed.
  • External exposure was not treated as current, owned, or vulnerable without proportionate authorized validation.
  • People-focused data is necessary, minimized, redacted, access-controlled, and safe to disseminate.
  • AI and automation outputs retain prompts/policy, tool calls, citations, validation, and reviewer decisions; no agent exceeded authority.
  • Confidence wording matches evidence quality, identity certainty, temporal relevance, source independence, and unresolved conflicts.
  • The handoff includes owners, actions, caveats, handling, monitoring or expiry, correction route, and retention/deletion.
  • A second analyst can reproduce at least one high-impact claim from the preserved record.

Current primary and practitioner references

Source set used to ground this guide

Use current official documentation and source-specific terms before operational collection. Product fields, quotas, visibility, APIs, and coverage change. The links below identify the maintained source rather than freezing every feature claim in this page.

OHCHR and UC Berkeley — Berkeley Protocol on Digital Open Source Investigations

Professional, legal, ethical, security, collection, preservation, verification, and reporting methodology for digital open-source investigations. Its formal human-rights context does not make every rule universal, but its evidence and safety discipline is broadly instructive.

MITRE ATT&CK — Reconnaissance tactic

Adversary reconnaissance behaviors including active scanning, victim organization/infrastructure/host/identity information, search of open websites and technical databases, phishing for information, and gathering victim network information.

NIST SP 800-115 — Technical Guide to Information Security Testing and Assessment

Planning, rules of engagement, test execution, analysis, reporting, and mitigation guidance. Apply it to active assessment under explicit authority, not as a blanket authorization.

OWASP Web Security Testing Guide

Maintained web-testing methodology with versioned information-gathering scenarios and reporting guidance. Distinguish stable/versioned material from current development.

ICANN — Registration Data Access Protocol

Current overview and resources for structured domain registration data access and RDAP’s relationship to legacy WHOIS.

RFC 9162 — Certificate Transparency Version 2.0

IETF Experimental RFC describing CT v2 public logging and auditing of TLS certificate issuance; it obsoletes RFC 6962 while CT ecosystem implementations may have version-specific realities.

OWASP Amass

Open-source attack-surface mapping and external asset discovery framework with collection, storage, and an Open Asset Model.

Shodan Help Center

Official source for banner semantics, query filters, CLI/API use, data timeframes, credits, monitoring, and scanning behavior.

Censys Platform documentation

Official current model for host, web-property, and certificate datasets, CenQL search, related assets, plan-dependent access, and direct lookups.

urlscan Documentation Hub

Official source for existing-scan search, API fields, visibility, submission, result retrieval, quotas, and data-source behavior.

VirusTotal API v3 concepts and relationships

Object, collection, relationship, pagination, and access semantics for files, URLs, domains, IPs, and related threat context.

CISA BOD 23-01 — Asset Visibility and Vulnerability Detection

FCEB-specific binding requirements and useful definitions distinguishing asset discovery from vulnerability enumeration. Do not present its cadence as a universal mandate.

1200km Network Reconnaissance library

First-party Nmap, CLI, Shodan, Censys, theHarvester, network discovery, and web-recon guides, with practitioner context and local article mirrors.

1200km Red Team reconnaissance and attack-surface module

Authorized active-testing context, asset modeling, safe validation, evidence, and reporting boundaries that complement this OSINT-focused guide.

AdversaryGraph · capabilities · IOC case study

1200km’s self-hosted analyst workbench and published workflow for source-backed IOC enrichment, typed pivots, ATT&CK leads, evidence review, and report handoff.