Orientation
Find less, understand more, and preserve the path between source and conclusion
OSINT quality is not measured by tabs opened, records collected, tools installed, or nodes drawn. It is measured by whether the collection answered a defined question, stayed inside authority and minimization boundaries, preserved provenance, separated observation from inference, survived challenge, and produced a proportionate decision.
Discovery
Identify candidate entities, infrastructure, aliases, services, publications, relationships, and changes from sources suited to the question. Record what each source actually observes and what it cannot establish.
Verification
Corroborate important claims through independent evidence, temporal consistency, source proximity, content analysis, and explicit alternatives. A repeated claim is not independent confirmation when every copy derives from one origin.
Protection
Protect investigators, subjects, sources, credentials, queries, workstations, case data, and downstream users. Minimize personal and sensitive data, control access, and avoid actions that create harm or tip off a monitored actor.
Communication
Deliver a claim-evidence matrix, timeline, entity model, confidence assessment, limitations, alternative hypotheses, and recommended collection or response. Make uncertainty visible instead of hiding it behind a polished graph.
Open-source information
Lawfully accessible material that may be observed, requested, purchased through legitimate channels, or obtained under an applicable public-access process. Accessibility, legality, terms, and ethical use remain context-dependent.
OSINT
An analyzed product derived from open-source information to address a defined requirement. Search output becomes intelligence only after evaluation, corroboration, interpretation, and communication.
Passive reconnaissance
Collection from existing third-party or supplied observations without intentionally sending assessment traffic to the target. “Passive” does not remove privacy, terms, attribution, or operational-security risks.
Active reconnaissance
Direct interaction with target-controlled systems or people to discover information. It requires explicit scope, rate and safety controls, source attribution, stop conditions, and coordination.
Operating model
One traceable chain from requirement to supported claim
Use a durable investigation chain: requirement → authority and risk assessment → hypotheses → collection plan → source capture → normalization → entity resolution → corroboration → analysis → confidence → review → dissemination → retention or deletion. Every transition should be inspectable.
| Layer | Question | Minimum record | Failure to avoid |
|---|---|---|---|
| Requirement | Which decision must this investigation support? | Requestor, decision, precise questions, priority, deadline, audience, acceptable uncertainty | Collecting everything about a person or organization “just in case” |
| Authority | What may be collected, from where, by which method, and for how long? | Purpose, legal/policy basis, scope, prohibited actions, privacy and safety constraints, approver | Assuming public visibility authorizes interaction, republication, or indefinite retention |
| Collection | Which source can observe the required fact? | Source, query, account/tool, timestamp, result, coverage, cost, terms, errors, capture hash | Treating an aggregator as the original source or a missing result as proof of absence |
| Identity | Do records refer to the same entity? | Stable identifiers, temporal overlap, shared attributes, conflicts, confidence, alternatives | Merging by name, email pattern, IP, logo, or shared hosting alone |
| Analysis | What is observed, inferred, assessed, and still unknown? | Claim-evidence matrix, hypotheses, confidence, source quality, contradictions, gaps | Graph proximity, AI prose, or repeated reporting presented as proof |
| Handoff | What action is proportionate and reviewable? | Finding, scope, owner, evidence links, caveats, sensitivity, next step, retention | Doxxing, over-attribution, automated enforcement, or publication without harm review |
Learning path
Fourteen-module OSINT and reconnaissance practitioner sequence
Complete the modules in order when building a new capability. Experienced researchers can use them as a review map, but the dependencies remain: collection without authority creates risk, data without provenance cannot be defended, and attribution without alternatives is storytelling.
Authority, ethics, privacy, safety, and operational security
Before searching, establish purpose, authority, scope, collection classes, prohibited methods, dissemination, retention, and stop conditions. Distinguish research about organization-controlled infrastructure from research about individuals. The latter can expose vulnerable people, families, employees, sources, investigators, or unrelated tenants, even when fragments are publicly visible.
Collection risk assessment
Assess legal and policy basis, platform terms, expected personal/special-category data, minors or vulnerable persons, source sensitivity, jurisdiction, necessity, proportionality, retention, reidentification risk, potential notification, and harm if the case file leaks. Record why each data class is needed.
Investigator OPSEC
Use organization-managed research identities where permitted, isolated browser profiles or workstations, separate credentials, strong MFA, controlled downloads, DNS and browser telemetry awareness, approved VPN/proxy design, malware-safe handling, logging, and a documented rule for when not to authenticate, click, download, contact, or join a community.
- Passive is not invisible: search providers, websites, social platforms, archives, and APIs can log queries, source addresses, cookies, browser characteristics, and account identity.
- Public is not ownerless: copyright, database rights, contractual terms, confidentiality, privacy, safeguarding, and contextual integrity may still apply.
- Contact changes the operation: messaging a subject, requesting access, creating a pretext, or joining a restricted group can create legal, ethical, evidentiary, and safety consequences.
- Data brokers are not neutral: provenance, consent, accuracy, access legitimacy, and downstream risk must be reviewed before acquisition or use.
Primary: OHCHR and UC Berkeley — Berkeley Protocol on Digital Open Source Investigations · NIST SP 800-115
Related Cyber Knowledge: Governance, Risk & Compliance (GRC) — Privacy engineering and data governance
Intelligence requirements, hypotheses, and collection planning
Translate a broad request into answerable questions. “Research this company,” “find everything about this person,” and “map the attack surface” have no natural stopping point. A defensible requirement names the decision, target entity, time window, relevant relationship, acceptable evidence, priority, constraints, and the uncertainty the decision-maker can tolerate.
Question decomposition
Break the requirement into claims: legal identity; owned domains and netblocks; externally operated services; brands and products; historical changes; supplier dependencies; public code; known malicious relationships; or whether one observable is connected to a campaign. State what would support, weaken, or falsify each hypothesis.
Collection matrix
For every question, list preferred original sources, secondary sources, query variants, identifiers, time coverage, expected result type, access and rate limits, privacy class, collection owner, validation route, fallback, and stop rule. Assign source priority before results bias the method.
Build the seed set without contaminating the conclusion
Start from known-good identifiers: official legal name, verified domain, registry identifier, supplied IP or hash, published account, or certificate fingerprint. Preserve where each seed came from. Expand one relationship at a time and label every edge by relationship type, observation time, and source. Never let a candidate discovered during collection silently become a confirmed seed.
1200km: Cyber Threat Intelligence field guide · AdversaryGraph IOC investigation workflow
Related Cyber Knowledge: Cyber Threat Intelligence (CTI) — Module 2 — The Intelligence Cycle Intelligence Types
Search engines, web archives, public records, and documents
General search is a discovery layer, not a source of truth. Query multiple indexes because coverage, ranking, geography, language, personalization, removals, caching, and date interpretation differ. Capture the result page, query, locale, time, and destination. Open the underlying document and locate its publisher, edition, date, context, and revision history.
- Query design: combine exact phrases, aliases, identifiers, site or domain restrictions, file types, date ranges, language variants, transliterations, known document titles, and exclusions. Keep a query notebook so another researcher can reproduce the path.
- Archives: compare captures over time, but distinguish capture time from publication time. Missing snapshots do not prove a page never existed. Archived scripts, redirects, images, robots exclusions, and dynamic content may be incomplete.
- Documents: inspect title, author, publisher, revision, embedded links, signatures, metadata, page references, and visual consistency. Sanitize downloads and avoid opening untrusted active content on normal workstations.
- Public records: use the authoritative registry, court, procurement, corporate, regulatory, standards, or government portal where available. Record jurisdiction, identifier, effective date, status, and whether the database is complete or delayed.
- Languages: preserve the original text and source. Treat machine translation as an aid, not a substitute for qualified review when nuance changes the conclusion.
1200km: Essential CLI tools for reconnaissance · Network reconnaissance guides
Related Cyber Knowledge: Cyber Threat Intelligence (CTI) — Module 4 — Collection Sources
Domain, DNS, RDAP, certificate transparency, and routing research
Infrastructure research works best as a time-aware evidence graph. A domain can resolve to shared CDN addresses; a certificate can cover customer-selected names; an ASN can announce provider and tenant space; and registration records can be privacy-protected or redacted. Every relationship needs a type, observation time, source, ownership interpretation, and confidence.
Names and registration
Use DNS for current records and authorized passive-DNS providers for historical observations. Use RDAP for structured domain, IP network, and autonomous-system registration data where supported. Preserve registrar/registry status, events, nameservers, entities returned, notices, redaction, and query time. Registration, administration, hosting, and operational control are different relationships.
Certificates and routing
Certificate Transparency logs can reveal names submitted in certificates and support monitoring for unexpected issuance. Inspect subject alternative names, issuer, validity, fingerprint, key and signature properties, source log, and first/last observation. Use RIR/RDAP and BGP sources for prefixes, origin AS, route history, and organization records while distinguishing allocation, announcement, hosting, and service ownership.
Resolution workflow
Start with the verified domain. Enumerate current authoritative DNS records and nameserver delegation. Search certificate records for exact and wildcard names, then validate candidates against current and historical DNS. Resolve IPs to registration and routing context. Identify CDN, cloud, shared-hosting, WAF, reverse proxy, email, SaaS, and third-party boundaries. A candidate becomes organization-attributed only when two or more suitable signals support that relationship or an authoritative source confirms it.
Primary: ICANN RDAP · RFC 9162 — Certificate Transparency Version 2.0 · 1200km: theHarvester guide
Related Cyber Knowledge: Red Team & Offensive Security — Module 2 — Reconnaissance and attack-surface mapping
Internet exposure, service indexes, and infrastructure search
Internet-wide search platforms expose observations collected at particular times through particular scan methods. Shodan’s fundamental record is a service banner; Censys models hosts, web properties, certificates, and related assets; urlscan stores website-scan observations and resources. Their outputs are valuable evidence leads, but they are not live ground truth, vulnerability proof, ownership proof, or permission to test.
Query discipline
Begin with confirmed organization identifiers: netblocks, exact domains, certificate names or fingerprints, approved organization fields, or known products. Use structured filters, narrow time windows, and exported raw records. Record account tier and coverage because history, fields, results, quotas, and API features vary.
Observation interpretation
Separate IP host, virtual host, web property, service, certificate, and screenshot observations. Note observed-at time, scanner vantage, protocol, port, transport, banner, HTTP/TLS metadata, labels, inferred software, CPE/CVE claims, and provider. Validate critical exposure against an authorized current source before escalation.
- Shodan: search banners and structured properties. Preserve the full banner and observation date; product or vulnerability fields remain scanner-derived claims.
- Censys: choose hosts for IP and non-HTTP service questions, web properties for hostname/HTTP behavior, and certificates for issuance and presentation relationships.
- urlscan: search historical scans before submitting a URL. Submission visibility matters; a public scan can disclose an investigation target or sensitive URL.
- Reputation platforms: distinguish crowdsourced labels, vendor verdicts, passive associations, detections, and analyst comments. Count sources, not repeated feeds that inherit one label.
Official: Shodan query fundamentals · Censys Platform quick start · urlscan Search API · 1200km: Shodan guide · Censys guide
Related Cyber Knowledge: Red Team & Offensive Security — Module 2 — Reconnaissance and attack-surface mapping
Web, API, mobile, and application reconnaissance
Application reconnaissance models entry points, trust boundaries, identities, data flows, technologies, deployments, and third parties. Passive discovery can include indexed pages, archived content, public documentation, application stores, package registries, public API specifications, DNS, certificates, client-side code already delivered to a normal user, and exposure-index observations. Direct crawling, directory enumeration, fingerprinting, parameter discovery, or protocol interaction is active testing and must stay inside explicit scope.
- Application inventory: hostname, scheme, ports, environment, product, owner, authentication, user roles, API base URLs, mobile package identifiers, status pages, documentation, repositories, data classes, and third-party integrations.
- Client-visible surface: links, scripts, source maps, API endpoints, feature flags, telemetry endpoints, storage names, CSP/connect destinations, public configuration, and framework metadata. Do not treat an endpoint name as proof that it is reachable or vulnerable.
- API context: published OpenAPI/GraphQL documentation, versioning, base paths, authentication scheme, rate-limit statements, SDKs, examples, deprecation, status, and public change logs.
- Historical surface: archived routes, old documentation, retired subdomains, mobile versions, abandoned marketing applications, and previous hosting. Validate current ownership and availability before classifying.
- Active phase: use an approved user agent, source address, rate, time window, request method set, wordlist, exclusion list, authentication account, data-handling rule, and emergency contact.
Primary: OWASP Web Security Testing Guide · 1200km: Web application reconnaissance · Secure Code / Application Security
Related Cyber Knowledge: Secure Code & Application Security — API, GraphQL, webhook, and business-logic security
People, organizations, brands, suppliers, and relationship research
People and organization research carries the highest risk of unnecessary collection and harm. Define the organizational question first: confirm an executive role from authoritative publications, identify the owner of a business service, validate a supplier relationship, detect brand impersonation, or understand public hiring signals relevant to a threat model. Avoid building personal dossiers where an organization-level answer is sufficient.
Authoritative organization sources
Prefer official corporate sites, filings, registries, regulator notices, procurement records, verified press releases, conference programs, patent or standards records, repositories, and the subject’s own professional publication. Record effective dates; titles, subsidiaries, brands, and supplier relationships change.
Identity resolution
Use multiple stable attributes: legal entity or employer, verified account, organization domain, location at the relevant time, publication history, unique identifier, or explicit cross-link. Names, profile photos, usernames, writing style, and contact patterns are weak alone. Preserve conflicts and avoid inferring sensitive attributes.
- Collect the minimum professional information necessary for the requirement and mask personal contact details in routine reporting.
- Distinguish current employee, former employee, contractor, speaker, contributor, customer, partner, reseller, and unverified self-claim.
- For brand monitoring, preserve the legitimate brand baseline—domains, certificates, visual marks, support channels, and campaigns—before labeling a lookalike.
- Do not contact colleagues, relatives, or unrelated third parties; do not attempt password recovery, account discovery, or identity challenge flows.
- Escalate safeguarding concerns and credible threats through approved channels rather than publishing personal information.
Primary: Berkeley Protocol · 1200km: Privacy / Data Handling · Privacy and data governance
Related Cyber Knowledge: Governance, Risk & Compliance (GRC) — Third-party, service-provider, and software supply-chain risk
Public code, packages, cloud artifacts, documents, and exposed secrets
Public repositories and package ecosystems can explain technologies, ownership, deployment patterns, identifiers, domains, storage names, API schemas, historical changes, and third-party dependencies. They can also expose credentials or sensitive configuration. Treat suspected secrets as hazardous evidence: do not test them, do not place them in tickets or chat, and do not broaden access.
- Repository discovery: search verified organizations, repository topics, package metadata, commit authorship where appropriate, documentation links, dependency files, CI workflows, infrastructure templates, container references, public issue trackers, and release artifacts.
- Attribution: a repository name, copied code, contributor email, package namespace, or cloud identifier may be stale, forked, mirrored, typosquatted, or unrelated. Confirm through organization ownership, signed releases, official links, registry publisher data, and temporal consistency.
- Secret handling: preserve only the minimum needed to report; mask values; record repository, path, commit, discovery time, type, exposure status, and responsible contact. Do not clone unnecessary history if a safer permalink or platform report is sufficient.
- Cloud artifacts: bucket names, tenant IDs, account IDs, regions, image names, workload identity, and endpoints are identifiers, not proof of public access. Validate ownership and exposure through the owner or an explicitly authorized assessment.
- Dependencies: package name and version may support technology inventory, but lockfiles, built artifacts, runtime inventory, and deployment evidence provide different confidence levels.
Official: GitHub repository visibility and security considerations · 1200km: Cloud Security · Secure Code / Application Security
Related Cyber Knowledge: Secure Code & Application Security — Dependencies, source control, builds, artifacts, and supply-chain assurance
Images, video, audio, geolocation, chronolocation, and media verification
Media verification asks whether an item is original or derivative, what it depicts, where and when it may have been captured, how it was published, and whether manipulation or context changes the claim. No single metadata field, reverse-image result, shadow estimate, AI detector, compression artifact, or visual resemblance should carry the conclusion.
Preserve before analysis
Capture the original URL, publisher/account, post and retrieval times, surrounding text, engagement context, available original file, hashes, file metadata, dimensions, codec/container, thumbnails, and archive reference. Work from copies and record transformations such as transcoding, frame extraction, cropping, enhancement, or translation.
Corroborate location and time
Compare landmarks, terrain, road geometry, signs, scripts, utilities, transit, vegetation, weather, sun direction, shadows, construction state, event schedules, satellite or map imagery, and independent media. Document which features are distinctive, which are common, and which alternatives remain possible.
- Use reverse-image and keyframe searches across multiple indexes to locate earlier publication and variants.
- Separate file creation metadata from upload, edit, transcoding, screenshot, and publication times.
- Do not expose precise locations of vulnerable people, shelters, medical facilities, witnesses, or ongoing operations.
- Track source lineage: original uploader, reposts, crops, translations, overlays, compilations, and news embeds can create false independence.
- Treat automated deepfake or manipulation scores as one signal with documented model/version/limitations, not a verdict.
Primary: Berkeley Protocol on Digital Open Source Investigations
Related Cyber Knowledge: Cyber Threat Intelligence (CTI) — Module 5 — Analysis Techniques Tradecraft
Threat infrastructure, IOC enrichment, campaign pivots, and attribution restraint
Threat-infrastructure research begins with a supplied observable or behavior and asks which relationships are supported: resolution, hosting, certificate presentation, content similarity, registration, malware communication, file download, redirect, tracker reuse, or co-observation in a report. Shared infrastructure and commodity services make careless clustering especially dangerous.
Observable discipline
Canonicalize IP, domain, URL, email, file hash, certificate fingerprint, mutex, registry path, user agent, or other artifact. Record source, first/last seen, role, confidence, TLP/handling, maliciousness basis, and expiration. Preserve exact URLs safely without browsing a live malicious destination from a normal workstation.
Relationship discipline
Use explicit predicates: domain resolved-to IP at time; certificate presented-by host; file contacted domain in sandbox run; URL redirected-to URL; report asserted actor-used infrastructure; registration shared email; content hash matched. “Related to” is too vague for analytic review.
Attribution ladder
Move from observation to cluster only when multiple independent or high-quality relationships align across time, infrastructure, behavior, malware, operations, and victimology. Keep infrastructure clustering, activity-cluster naming, malware-family assessment, campaign assessment, and sponsor attribution as separate conclusions. A TTP overlap, shared hosting provider, common registrar, commodity tool, or AI-generated similarity score is a lead—not attribution.
Official: VirusTotal relationship model · MITRE ATT&CK Reconnaissance tactic · 1200km: From log to report with AdversaryGraph · AdversaryGraph IOC investigation capabilities
Related Cyber Knowledge: Cyber Threat Intelligence (CTI) — Module 6 — The Threat Actor Landscape
Automation, APIs, entity graphs, normalization, and data engineering
Automation should make collection reproducible and reviewable, not turn every weak association into a permanent fact. Design pipelines around immutable raw capture, normalized records, typed relationships, provenance, temporal validity, confidence, access control, review state, and deletion. Preserve the ability to rebuild derived graphs when parsing or resolution logic changes.
- Raw layer: source response, request metadata, collection timestamp, tool and version, account tier, status/error, response headers where relevant, content hash, encryption, and retention class.
- Normalized layer: canonical entity type and value, stable internal ID, source ID, observation time, valid-from/to, labels, sensitivity, and parser version. Keep raw values for audit.
- Relationship layer: subject, predicate, object, source, observation time, confidence, inference rule, reviewer, status, and expiry. Avoid generic edges.
- Analytic layer: hypotheses, claim-evidence links, source independence, scoring explanation, contradictions, gaps, and review history.
- Operational layer: case, owner, dissemination, alert/detection/remediation links, suppression, retention, and deletion proof.
Reliability engineering
Honor rate limits, quotas, acceptable-use terms, pagination, time zones, retries, backoff, schema changes, partial responses, and API deprecation. Cache according to source terms and freshness needs. Deduplicate by source-native ID plus canonical observation—not by value alone—so multiple independent sources and historical changes remain visible. Quarantine malformed or ambiguous records instead of silently coercing them.
Official tooling: OWASP Amass · urlscan search semantics · 1200km: AdversaryGraph platform guide
Related Cyber Knowledge: Cyber Threat Intelligence (CTI) — Module 9 — Tools of the Trade
AI-assisted OSINT with bounded tools, citations, and human review
Apply the AI Security agent and MCP control model before connecting a model to search, enrichment, browser, scanning, or publication tools.
AI can accelerate query generation, translation, document triage, entity extraction, schema normalization, relationship suggestions, timeline drafting, hypothesis generation, and report formatting. It can also fabricate sources, merge identities, expose sensitive case data, follow malicious content instructions, overstate confidence, and initiate actions outside scope when connected to tools.
Safe task design
Provide a structured requirement, approved source set, data classification, target schema, prohibited inferences, citation requirement, confidence vocabulary, and instruction to return unknown when evidence is absent. Use retrieval over approved case material and require source spans for every proposed claim.
Tool and MCP controls
Expose read-only, allowlisted tools by default. Separate search from active scanning, authentication, messaging, purchase, upload, and publication. Require explicit human authorization for any state-changing or target-interacting action. Log tool name, arguments, source, response hash, model, prompt version, and reviewer decision.
- Redact unnecessary personal data, secrets, private sources, and customer information before remote processing; follow provider data-handling policy.
- Treat webpages, documents, repository content, images, and search results as untrusted data that may contain prompt-injection instructions.
- Resolve entities deterministically where possible. AI may suggest a merge, but a reviewer must inspect supporting and conflicting identifiers.
- Recalculate dates, IP ranges, hashes, DNS names, geographic coordinates, and query syntax with deterministic validation.
- Preserve citations to the original item, not to the model response or a generated summary.
- Evaluate on known cases for source fidelity, identity precision/recall, abstention, privacy leakage, unsafe tool selection, and confidence calibration.
1200km: AI & Offensive Security research · HexStrike laboratory index · Shodan with HexStrike AI · AI-assisted OSINT exposure-map case study
Related Cyber Knowledge: AI Security — AI red teaming, evaluation, and reproducible security testing
Defensive external attack-surface discovery and continuous monitoring
A defensive exposure program reconciles what the organization believes it owns with what external observers can see. Its goal is not a one-time internet scan. It maintains an attributed inventory of domains, certificates, IPs, netblocks, cloud and SaaS properties, applications, APIs, repositories, mobile apps, brands, and suppliers; detects meaningful change; verifies ownership; routes remediation; and measures closure.
Discovery-to-inventory loop
Seed from legal entities, brands, known domains, cloud accounts, certificates, netblocks, repositories, and service owners. Discover candidates through approved sources. Score ownership evidence separately from exposure risk. Ask the responsible owner to confirm or reject. Add confirmed assets with lifecycle, criticality, environment, dependency, and monitoring coverage.
Change-to-action loop
Track first/last seen, new certificate, DNS change, new port/service, login or administration page, exposed development environment, new repository/package, storage reference, security-header change, or brand lookalike. Deduplicate, suppress expected provider churn, verify current state, and route by owner and severity.
Metrics that support decisions
Measure candidate-to-confirmed precision, unknown-owner backlog, time to ownership decision, time to verification, exposure age, stale assets, orphaned domains, unexpected certificates, high-risk changes, remediation age, false-positive causes, source freshness, and coverage by critical service. Avoid using the raw number of discovered assets or scanner “risk score” as a maturity grade.
Primary: CISA BOD 23-01 (mandatory only for its stated FCEB scope; useful visibility concepts elsewhere) · 1200km: AdversaryGraph Asset Surface guide · Cloud posture and exposure
Related Cyber Knowledge: Cloud Security — Posture management, exposure, vulnerabilities, attack paths, and validation
Analysis, confidence, evidence preservation, reporting, and operational handoff
The final product should make the reasoning inspectable. Separate facts observed directly, source assertions, analytic inferences, and unknowns. Use consistent confidence language tied to evidence quality, source independence, identity resolution, temporal relevance, contradictions, and alternative explanations. Confidence is not the same as impact, priority, or probability.
Claim-evidence matrix
For each claim, record supporting observations, contradicting observations, source proximity, independence, timestamps, identity rationale, assumptions, alternative hypotheses, confidence, and reviewer. Link graphs, timelines, and screenshots back to the matrix rather than letting visuals become unsupported conclusions.
Preservation package
Maintain original capture, normalized record, query and collection log, hashes, timestamps, tool/version, transformations, access and handling, analyst notes, source restrictions, review history, and export. For legal or disciplinary use, obtain qualified evidence-handling guidance and preserve chain-of-custody requirements.
- Executive output: question, supported answer, confidence, material implications, limitations, decision options, and next review.
- Operational output: confirmed entities/observables, scope, owner, timeline, typed relationships, source links, detection/remediation actions, expiry, and handling.
- Analytic annex: methodology, sources, queries, identity resolution, alternatives, gaps, contradictions, AI/tool use, and quality controls.
- Redaction: remove personal contact data, credentials, irrelevant private details, precise safeguarding-sensitive locations, and provider secrets.
- Correction: provide a route to challenge or correct a record; preserve what changed, why, by whom, and which downstream products were updated.
1200km: External validation and evidence · DFIR evidence and timeline practice · AdversaryGraph campaign-cluster investigation
Related Cyber Knowledge: Cyber Threat Intelligence (CTI) — Module 7 — Intelligence Products Sharing
Source strategy
Choose sources by observation model, not brand familiarity
No platform “has all the data.” Select a source because it can observe the entity or relationship needed, then document coverage, timing, terms, and validation. The matrix below is a decision aid, not an endorsement or claim of completeness.
| Question | Preferred source class | What it can support | Essential caveat |
|---|---|---|---|
| Who is the registered domain/IP/ASN entity? | Registry/RIR RDAP and authoritative registration services | Structured registration object, status, events, entities returned, notices | Redaction and privacy proxies are common; registration is not operational ownership |
| What does a name resolve to now or historically? | Authoritative/current DNS and approved passive-DNS providers | Dated record observation and resolution history | Resolver cache, split horizon, sinkholes, wildcarding, and provider coverage affect results |
| Which names appeared on certificates? | CT logs, certificate datasets, live TLS where authorized | Certificate issuance/logging and presentation observations | A certificate name or wildcard does not prove live service or organization ownership |
| Which public services were observed? | Internet-wide host/web-property search platforms | Dated banner, service, protocol, web, TLS, or screenshot observations | Records may be stale, shared, parsed incorrectly, or different from current state |
| What pages/resources were loaded? | Website-scan archives and owner-authorized browsing/capture | Dated page, request, response, resource, redirect, and visual context | Submitting a new scan may interact with or publicly disclose the target |
| What did the organization publish? | Official site, filings, registry, release, documentation, repository | First-party assertion at a time and version | Self-published claims still require interpretation and may later change |
| How is an observable related to malware? | Sandbox/report evidence and typed threat-intelligence relationships | Observed communication, download, execution, or asserted report relationship | Reputation aggregation and shared infrastructure can produce misleading associations |
| What existed in the past? | Web archives, historical DNS/cert/service records, version control | Point-in-time capture or change history | Archive coverage is incomplete; absence of capture is not absence of existence |
| Does an asset currently belong to us? | Authoritative inventory, account/provider data, owner confirmation | Current accountable ownership and lifecycle | External observation alone rarely establishes full ownership |
Reusable artifacts
Templates that preserve method, evidence, and restraint
A tool output is not an investigation record. These artifacts connect collection to decision and make the work reviewable, correctable, and repeatable.
Investigation charter
Case ID; requestor; decision; priority questions; target entity and known-good seeds; purpose and authority; jurisdictions; approved source/method classes; prohibited actions; personal/sensitive data; safeguarding; research identity; active-contact rule; dissemination; retention; owner; approver; start/stop conditions.
Collection record
Collection ID; question; source and dataset; exact query/request; account/tool/version; collection and observation time; locale/time zone; result and pagination; raw hash; source terms/classification; errors/coverage; extracted entities; analyst; review status.
Entity record
Internal ID; entity type; canonical value/name; aliases; stable external identifiers; source observations; valid-from/to; ownership/role; sensitivity; candidate/confirmed/rejected state; merge and split rationale; conflicts; confidence; reviewer; expiry.
Typed relationship record
Subject; predicate; object; direct observation or inference; source; observed-at; valid time; first/last seen; supporting and contradicting evidence; shared-service context; inference rule; confidence; analyst/reviewer; expiry; downstream use.
Claim-evidence matrix
Claim; decision relevance; supporting evidence; contradictions; source proximity and independence; temporal fit; identity fit; alternative hypotheses; assumptions; gaps; confidence; wording; reviewer; correction status.
Exposure candidate
Candidate asset; discovery source; ownership signals; provider/tenant boundary; environment; service; current verification; risk lead; owner; confirmation status; first/last seen; remediation; closure evidence; inventory and monitoring update.
AI-assistance audit
Task; provider/model; prompt/policy; retrieval set; redaction; tool calls; output schema; citations; validation; unsafe or rejected suggestions; reviewer changes; accepted claims; cost/latency where material; evaluation version.
Case preservation manifest
Original files/URLs; capture method; hashes; timestamps; transformations; storage; access history; handling; source restrictions; derived files; reports; dissemination; legal hold where authorized; correction; retention trigger; deletion proof.
End-to-end practice
Six OSINT and reconnaissance case studies
These cases are designed for owned infrastructure, supplied datasets, documentation, or approved lab material. They demonstrate method and evidence handling without encouraging unscoped targeting of real people or third-party systems.
Case 1 — Reconcile a company’s external web estate
- Create known-good seeds from legal and inventory records; do not begin from an internet-wide fuzzy organization search.
- Collect DNS, RDAP, CT, approved passive DNS, host/web-property, repository, status-page, and application-store observations.
- Normalize FQDN, certificate, IP, ASN, repository, application, and provider entities; retain observation times.
- Classify direct ownership, acquired entity, SaaS, CDN/WAF, supplier, shared hosting, historical, candidate, or rejected.
- Ask owners to confirm candidates and environment; compare with cloud/account inventory and service maps.
- Verify material exposure through approved current sources; create owner-bound remediation and monitoring records.
Case 2 — Investigate a suspected phishing domain
- Canonicalize the indicator and preserve the reporting source without opening it from a normal browser.
- Search existing URL-scan, DNS, RDAP, CT, reputation, certificate, redirect, screenshot, and historical observations before any new submission.
- Compare string, visual, certificate, infrastructure, content, and timing with the legitimate brand baseline.
- Identify shared hosting/CDN, registrar, nameserver, campaign-related infrastructure, and malware/report evidence without equating provider reuse with actor identity.
- Assign role and confidence; route domain/URL and supporting evidence to blocking, takedown, fraud, legal, or monitoring owners.
- Set expiry and recheck; record sinkhole, takedown, content, DNS, and certificate changes.
Case 3 — Validate an unexpected certificate alert
- Preserve certificate fingerprint, SANs, issuer, validity, CT log/source, and alert time.
- Classify exact, wildcard, lookalike, internal-looking, customer-controlled, provider, or unrelated names.
- Query DNS and authorized historical sources; compare known certificate-management providers and account records.
- Contact the named certificate/domain owner through internal channels; do not interact with an unknown endpoint solely because it appeared in CT.
- Determine authorized issuance, abandoned DNS, provider validation artifact, brand abuse, or false candidate.
- Update inventory, issuance monitoring, CAA/control review, response, and closure evidence.
Case 4 — Triage a public-repository secret report
- Restrict access to the report; preserve repository, commit, path, time, secret type, and a masked excerpt.
- Confirm repository and organization relationship through authoritative links; identify fork, mirror, archived, or unrelated status.
- Notify the security/secret owner through approved private channels and revoke/rotate before broader investigation.
- Use platform history and internal provider logs—not authentication attempts—to determine validity and use.
- Remove or rewrite history where appropriate, enable detection and push protection, and inspect related exposure under incident authority.
- Delete unnecessary copies; document revocation, remediation, monitoring, and disclosure handling.
Case 5 — Build a threat-infrastructure cluster from one IOC
- Validate type, source, role, first/last seen, handling, and why the seed is malicious or suspicious.
- Collect direct DNS, certificate, hosting, URL, redirect, sandbox, file, report, and registration relationships.
- Normalize time and separate shared, benign, scanner, sinkhole, CDN, and unknown infrastructure.
- Build candidate clusters using documented rules; require multiple high-quality signals for promotion.
- Compare behaviors and report evidence; map ATT&CK only from described behavior, not infrastructure type.
- Create detection/blocking/hunting leads with confidence and expiry; retain alternatives and contradictions.
Case 6 — AI-assisted review of a supplied research package
- Classify and redact the package; create immutable originals and a retrieval corpus limited to approved evidence.
- Define schemas for entities, observations, typed relationships, claims, citations, contradictions, and unknowns.
- Ask the model to extract and propose—not decide—while treating source content as untrusted data.
- Validate identifiers, dates, citations, translations, calculations, and entity merges deterministically.
- Have an analyst review every claim and compare alternative hypotheses; reject uncited or overconfident output.
- Save provider/model/prompt/retrieval/tool/reviewer audit and generate the report from accepted structured records.
Hands-on curriculum
Twelve-lab OSINT and reconnaissance sequence
Use your own domain, an intentionally published lab domain, reserved examples such as example.com, supplied offline datasets, and platform documentation. Do not redirect these exercises at an arbitrary real person or organization.
Lab 01 — Investigation charter
Turn “map our exposure” into three precise questions. Write scope, authority, source classes, prohibited methods, privacy risks, research identity, active-contact rule, retention, completion criteria, and escalation. Peer-review the charter before collecting.
Lab 02 — Query and provenance notebook
Use a harmless subject to run exact-phrase, site, file-type, date, language, alias, and identifier queries across two search engines. Preserve query, locale, time, result, original source, archive state, and what each missing result cannot prove.
Lab 03 — DNS/RDAP/CT evidence graph
For a domain you control, collect current DNS, ICANN/RIR RDAP, CT records, live certificate metadata, and provider context. Build typed dated edges and mark authoritative ownership, service provider, candidate, and historical relationships separately.
Lab 04 — Internet-index comparison
Compare approved Shodan, Censys, and urlscan observations for your own asset. Record dataset, query, observation time, host versus web-property semantics, stale/shared results, and current verification. Explain every disagreement.
Lab 05 — Passive application map
Map your own public website from normal browser-delivered HTML, scripts, CSP destinations, documentation, status pages, archives, DNS, and certificates. Separate observed endpoints from inferred or historical candidates; perform no fuzzing.
Lab 06 — Repository and package inventory
Use a repository you own. Trace organization, packages, releases, dependency files, container images, CI workflows, domains, and cloud identifiers. Create a private synthetic secret example and practice masked reporting without using the value.
Lab 07 — Media verification worksheet
Use an instructor-supplied non-sensitive image and known answer. Preserve file/context, extract metadata and frames, identify distinctive features, search variants, test two location/time hypotheses, and document confidence and transformation history.
Lab 08 — IOC relationship graph
Use a historical, defanged IOC package supplied with source records. Normalize entities; type DNS, certificate, URL, file, and report edges; identify shared infrastructure; create candidate clusters; add expiry; and write why the graph does not prove attribution.
Lab 09 — Source/API ingestion
Ingest a small saved JSON response. Keep raw hash and request metadata, validate schema, normalize observations, preserve source IDs, quarantine malformed rows, deduplicate without losing independent sources, and rebuild the graph from raw data.
Lab 10 — AI extraction evaluation
Provide an approved offline report with known entities and citations. Evaluate extraction precision/recall, identity merges, citation resolution, abstention, prompt-injection resistance, sensitive-data handling, and analyst correction. No external tools are enabled.
Lab 11 — External asset reconciliation
Combine a synthetic CMDB with DNS, certificate, and service observations containing shared-provider traps. Score ownership evidence, request simulated owner dispositions, confirm inventory changes, route exposures, and measure candidate precision and unknown-owner backlog.
Lab 12 — Defensible final product
Produce a charter, source/query register, entity dictionary, typed graph, timeline, claim-evidence matrix, confidence assessment, alternatives, redacted executive summary, operational handoff, correction route, and retention action. A second analyst must reproduce one claim.
Failure atlas
Common ways OSINT becomes unsafe or analytically weak
- Collection without a decision“Find everything” creates uncontrolled scope, privacy risk, and a pile of facts no one can use. Rewrite the requirement and define stopping conditions.
- Public-equals-permitted thinkingVisibility does not authorize active probing, contact, republication, or indefinite retention. Recheck authority, terms, necessity, and harm.
- Name-only identity mergeCommon names and reused usernames create false people and organizations. Require stable corroborating identifiers and preserve alternatives.
- Shared-infrastructure ownershipA CDN IP, wildcard certificate, cloud ASN, nameserver, or analytics ID pulls unrelated tenants into the graph. Model provider relationships explicitly.
- Timestamp collapseCollection, publication, observation, registration, certificate validity, archive capture, and event time become one “date.” Preserve each semantic time separately.
- Aggregator circularitySeveral platforms repeat one original report, creating apparent corroboration. Trace lineage and count independent origins.
- Absence as proofNo search result becomes “does not exist.” State source set, window, coverage, errors, and “not observed.”
- Stale exposure escalationAn old banner or screenshot becomes an urgent current vulnerability. Verify ownership and current state through an authorized source.
- Credential verificationAn investigator tests a discovered secret. Revoke and investigate through authorized logs; never authenticate to “confirm.”
- Public URL submissionA sensitive target is submitted to a public scanner, exposing the investigation. Search existing records first and select visibility deliberately.
- Graph-as-proofLayout, color, centrality, or node proximity replaces typed evidence. Make every edge reviewable and keep candidate status visible.
- AI citation launderingGenerated prose cites a search snippet or nonexistent source. Resolve every citation to original captured evidence before acceptance.
- Unbounded agent toolsAn AI system expands scope, scans, contacts, or publishes. Enforce allowlists, read-only defaults, approval gates, logging, and stop controls.
- Investigator exposureNormal accounts, browser profiles, or workstations reveal identity or ingest malicious content. Use approved research identities and isolated handling.
- Overcollection of peoplePersonal details enter a company-security case without necessity. Minimize, redact, restrict, and delete; escalate safeguarding issues.
- No correction pathA false association persists across graphs, reports, alerts, and blocklists. Version claims and propagate correction or deletion downstream.
Readiness gate
Accept an OSINT product only when the method and limits are visible
- The product names the decision, requestor, questions, scope, time window, and completion criteria.
- Authority, privacy, platform terms, investigator safety, source handling, contact rules, and retention were reviewed.
- Known-good seeds and candidate discoveries remain distinguishable throughout the case.
- Every material observation records source, query or retrieval route, collection time, observation/publication time where available, and coverage caveat.
- Original captures are hashed or otherwise integrity-protected; transformations and tools are documented.
- Entities were resolved using appropriate stable identifiers; weak name, IP, certificate, or hosting matches were not silently merged.
- Relationships use explicit predicates, evidence, temporal bounds, confidence, and expiry rather than generic “related” edges.
- Original, secondary, and aggregate sources are distinguished; apparent corroboration was checked for common lineage.
- Facts, source assertions, inferences, assumptions, and unknowns are visibly separated.
- Alternative hypotheses and contradicting evidence were recorded and reviewed.
- External exposure was not treated as current, owned, or vulnerable without proportionate authorized validation.
- People-focused data is necessary, minimized, redacted, access-controlled, and safe to disseminate.
- AI and automation outputs retain prompts/policy, tool calls, citations, validation, and reviewer decisions; no agent exceeded authority.
- Confidence wording matches evidence quality, identity certainty, temporal relevance, source independence, and unresolved conflicts.
- The handoff includes owners, actions, caveats, handling, monitoring or expiry, correction route, and retention/deletion.
- A second analyst can reproduce at least one high-impact claim from the preserved record.
Current primary and practitioner references
Source set used to ground this guide
Use current official documentation and source-specific terms before operational collection. Product fields, quotas, visibility, APIs, and coverage change. The links below identify the maintained source rather than freezing every feature claim in this page.
OHCHR and UC Berkeley — Berkeley Protocol on Digital Open Source Investigations
Professional, legal, ethical, security, collection, preservation, verification, and reporting methodology for digital open-source investigations. Its formal human-rights context does not make every rule universal, but its evidence and safety discipline is broadly instructive.
MITRE ATT&CK — Reconnaissance tactic
Adversary reconnaissance behaviors including active scanning, victim organization/infrastructure/host/identity information, search of open websites and technical databases, phishing for information, and gathering victim network information.
NIST SP 800-115 — Technical Guide to Information Security Testing and Assessment
Planning, rules of engagement, test execution, analysis, reporting, and mitigation guidance. Apply it to active assessment under explicit authority, not as a blanket authorization.
OWASP Web Security Testing Guide
Maintained web-testing methodology with versioned information-gathering scenarios and reporting guidance. Distinguish stable/versioned material from current development.
ICANN — Registration Data Access Protocol
Current overview and resources for structured domain registration data access and RDAP’s relationship to legacy WHOIS.
RFC 9162 — Certificate Transparency Version 2.0
IETF Experimental RFC describing CT v2 public logging and auditing of TLS certificate issuance; it obsoletes RFC 6962 while CT ecosystem implementations may have version-specific realities.
OWASP Amass
Open-source attack-surface mapping and external asset discovery framework with collection, storage, and an Open Asset Model.
Shodan Help Center
Official source for banner semantics, query filters, CLI/API use, data timeframes, credits, monitoring, and scanning behavior.
Censys Platform documentation
Official current model for host, web-property, and certificate datasets, CenQL search, related assets, plan-dependent access, and direct lookups.
urlscan Documentation Hub
Official source for existing-scan search, API fields, visibility, submission, result retrieval, quotas, and data-source behavior.
VirusTotal API v3 concepts and relationships
Object, collection, relationship, pagination, and access semantics for files, URLs, domains, IPs, and related threat context.
CISA BOD 23-01 — Asset Visibility and Vulnerability Detection
FCEB-specific binding requirements and useful definitions distinguishing asset discovery from vulnerability enumeration. Do not present its cadence as a universal mandate.
1200km Network Reconnaissance library
First-party Nmap, CLI, Shodan, Censys, theHarvester, network discovery, and web-recon guides, with practitioner context and local article mirrors.
1200km Red Team reconnaissance and attack-surface module
Authorized active-testing context, asset modeling, safe validation, evidence, and reporting boundaries that complement this OSINT-focused guide.
AdversaryGraph · capabilities · IOC case study
1200km’s self-hosted analyst workbench and published workflow for source-backed IOC enrichment, typed pivots, ATT&CK leads, evidence review, and report handoff.