Prospect Intelligence Platform
A safety-first Phase 8 design/implementation boundary for manual, evidence-led prospect qualification, controlled source ingestion, bounded website scanning, and domain intelligence review. Phase 8 website scanning is a conservative observation workflow: it never submits forms, executes JavaScript, follows unsafe protocols, or authorizes outreach. Automated outreach is disabled, and no live source may be enabled without explicit approval.
Included
- Dependency-free Python/SQLite API under
apps/api. - Tenant-scoped business detail APIs with child intelligence/evidence records, provenance fields, notes, pipeline state, and audit history.
- Server-side normalization, conservative website classification, exact deduplication, versioned scoring, and suppression checks. Phase 6 documents the SA phone/location canonical forms and the review-only fuzzy-match contract.
- Bounded list pagination and server-side filters so a tenant cannot request an unbounded prospect collection.
- Responsive static dashboard under
apps/webwith authenticated explorer filters, paginated results, detail review, manual intake, notes/pipeline context, evidence provenance, and browser-only CSV preview. - Docker Compose runtime with non-root containers, read-only filesystems, health checks, and a named SQLite data volume.
- Browser authentication with server-side sessions and an optional first-run admin bootstrap.
- Phase 4 MVP job monitor and SQLite-backed job/event schema/API surface, with the production limitations documented below.
- Phase 5 source-ingestion contract: an approved source registry owns adapter terms, rate limits, retention, and health/circuit policy; CSV and manual reference adapters are the safe initial adapters.
- Discovery queries are recorded as bounded, auditable intent and dry-run plans. Recording a query does not perform network discovery or imply that results exist.
Current workflow and Phase 4/5 boundary
- A permitted workspace member manually creates or reviews a prospect.
- The business detail response is the aggregate record for that tenant; related intelligence/evidence rows are returned only through the tenant-scoped detail surface.
- Each manually entered intelligence item should retain its source/provenance (for example, source label or URL, observed value, and captured/verified time). Missing provenance is a data-quality limitation, not permission to infer facts.
- Members use the pipeline state and notes to coordinate human review. A state change or note is an application event and is included in the record's audit/activity history where exposed by the API.
- Suppression remains a hard safety boundary. Suppressed or unreviewed records must not be treated as eligible for contact.
The API applies the organization/tenant boundary server-side to list, detail, child-record, notes, pipeline, and audit reads and writes. Clients must use the returned pagination metadata and follow next/previous links or tokens rather than assuming that one response contains the whole tenant dataset. See apps/api/README.md for the route contract and limits.
Phase 4 jobs/live logging contract
The planned asynchronous contract is: create one tenant-scoped job, return a stable job identifier, and move it through queued → running → a terminal state (succeeded, failed, cancelled). Each accepted request should carry an idempotency key whose scope and request fingerprint prevent duplicate jobs while allowing a safe replay of the original result. A job should persist append-only events with a monotonically increasing per-job sequence number, timestamp, level/type, safe message, and job/tenant identifiers.
Clients should poll a tenant-scoped job status/events endpoint using after_sequence (or an equivalent cursor), with bounded backoff and terminal-state handling. SSE is a planned low-latency delivery option, not a current implementation; polling remains the compatibility fallback. Cancellation and retry must be explicit, authorized controls: cancellation is cooperative and may finish as cancelled or report that the job is already terminal; retry creates a new attempt while retaining the original job/idempotency lineage and must not duplicate side effects.
The current MVP has SQLite job/event persistence, job status/list/detail and event APIs, cancellation/retry controls, and a browser job monitor that polls while work is active. SSE is not implemented; it remains a future delivery optimization over the persisted cursor. There is no Redis/Celery worker: the current in-process worker is suitable only for development/pilot use and must not be treated as durable, horizontally scalable execution.
Run locally
cd apps/api
python3 -m unittest discover -v
python3 app/main.py --host 127.0.0.1 --port 8000 --db /tmp/prospects.db
Serve the UI separately:
cd apps/web
python3 -m http.server 8080
Open http://127.0.0.1:8080. Set window.API_BASE in the browser console to http://127.0.0.1:8000 when testing the authenticated API locally, then sign in with the configured workspace credentials.
API smoke calls
curl http://127.0.0.1:8000/api/v1/health/live
curl 'http://127.0.0.1:8000/api/v1/businesses?page=1&page_size=25&pipeline_stage=new'
curl http://127.0.0.1:8000/api/v1/businesses/1
curl -X POST http://127.0.0.1:8000/api/v1/businesses \
-H 'content-type: application/json' \
-d '{"name":"Example Plumbing","website":"https://example.invalid","email":"info@example.invalid","phone":"+27 21 555 0100"}'
The protected calls require the authenticated session cookie. Exact child-record, notes, pipeline, and audit routes are documented in apps/api/README.md and are never cross-tenant addressable by changing an ID.
Compose
cp .env.example .env
docker compose config --quiet
docker compose up --build -d
curl -fsS http://localhost:8000/api/v1/health/live
curl -fsS http://localhost:8080/healthz
docker compose down
Compose passes the optional BOOTSTRAP_ADMIN_EMAIL and BOOTSTRAP_ADMIN_PASSWORD values to the API. Set both in an untracked .env only when provisioning a fresh instance, then remove them and rotate the password after the bootstrap admin is created. No credentials belong in this repository.
Authenticated browser requests use a server-side session cookie; login creates a session and logout invalidates it. The liveness endpoints (GET /api/v1/health/live and GET /healthz) intentionally remain unauthenticated so Docker, ingress, and monitoring health checks can use them. Authentication is not a substitute for tenant/authorization checks: protected routes must enforce the session and organization boundary server-side.
Phase 5 source boundary and remaining limitations
Phase 5 defines a source adapter contract and registry; Phase 8 adds a bounded website-observation adapter, but it does not implement general network discovery, enrichment scheduling, or a live external-source adapter. A source adapter must declare its identity, terms owner, permitted purpose, rate limits, retention class, query/result schema, dry-run behavior, and health/circuit controls. CSV and manual reference adapters may be used for operator-supplied data; they must preserve source attribution and raw source records, and must not silently turn preview data into outreach or verified facts.
A discovery query is a tenant-scoped, bounded, auditable request that can be validated and dry-run without contacting a source. Any live source requires explicit product/legal/security approval, a registered adapter, and an operational enablement decision; absent all three, execution must fail closed. Circuit-open, rate-limit, terms, or approval failures must produce a safe non-live result. Raw source records are retained only under the approved retention class and must exclude secrets and unnecessary personal data.
SQLite, the in-process worker, and the named local volume are suitable for the pilot only; production migration, durable queue/worker leases, event retention/backup, SSE delivery, and tested backup/restore remain unfinished. Redis and Celery are not implemented. The development password fallback is PBKDF2 rather than production Argon2id. Before production, complete the gates in docs/SECURITY.md and docs/OPERATIONS.md, including MFA, TLS, CSRF protection, source approval and terms review, rate limiting, circuit monitoring, tenant-scoped job/event authorization, idempotent side-effect handling, durable raw-source/audit retention, SSRF-safe fetching if a future scanner is approved, and tested backups/restores.
Phase 6 normalization and deduplication boundary
Normalization is deterministic and versioned. For South African data, phone values are stripped to digits, local 10-digit 0 forms and 00 27 forms are converted to canonical +27..., and unknown international numbers retain their explicit country code; presentation punctuation must not create a second identity. Locations derive whitespace/case/diacritic-folded province, city, and suburb fields. A normalized value is not proof that the underlying observation is correct.
Exact keys (for example, canonical domain, email, or phone) may identify duplicate candidates. Fuzzy matching is deterministic and suggestion-only: the same inputs and normalization version produce the same candidate, score, and reason. A suggested match must never merge automatically. Use the documented thresholds: >=0.90 is a strong suggestion, 0.75–0.8999 is a review suggestion, and <0.75 is not surfaced as a suggestion. A human with permission must explicitly confirm each merge.
Every confirmed merge must create a tenant-scoped, immutable-enough merge snapshot before mutation, recording the surviving and absorbed IDs, normalized comparison inputs, score/reasons, acting user, timestamp, and schema/normalization versions. The operation must be reversible from that snapshot. It must preserve or re-parent every child, evidence item, provenance/source-record link, note, pipeline/audit history, and original source identity; conflicts remain visible for human resolution rather than being silently overwritten. Cross-tenant candidates are never comparable or mergeable, and each suggestion, confirmation, rejection, reversal, and preservation/conflict decision belongs in the audit trail.
The MVP now exposes deterministic match suggestions at GET /api/v1/businesses/{id}/matches, explicit merge confirmation in the web review dialog, tenant-scoped merge history, and POST /api/v1/merge-history/{id}/reverse. The implementation remains a pilot boundary: hardening is still needed for a dedicated merge permission, stronger server-side confirmation semantics, full snapshot conflict handling, and production-grade rollback guarantees. Do not describe a normalized or suggested match as verified identity, discovery, enrichment, or outreach authorization.
Phase 7 domain intelligence boundary
Domain intelligence is an observation and review aid, not proof of business identity, control of a domain, or availability. A domain normalizer may derive a lowercase ASCII/Unicode comparison form and a registrable domain using a versioned Public Suffix List (PSL). The PSL is an input with update/version drift: unknown, private, malformed, single-label, localhost, and IP-literal values must remain unresolved rather than guessed. A subdomain is not automatically a separate candidate, and a public suffix itself is never a registrable domain.
DNS status is explicit: not_checked, pending, resolved, nxdomain, no_data, timeout, servfail, blocked, and error describe the check outcome, not a business conclusion. MX, NS, and TXT observations may be absent, partial, truncated, stale, resolver-dependent, or blocked; no record is not proof that mail, delegation, ownership, or a business relationship is absent. Store the resolver/source, observed time, TTL where supplied, and uncertainty/error metadata. Caches must be bounded and keyed by normalized query/type/class plus resolver policy, honor an observed TTL without extending authority, and expose freshness/staleness; cached data must never be presented as a fresh check.
Association confidence is separate from DNS status and from duplicate score. It must be derived from explainable, tenant-scoped evidence (for example, an operator citation, an exact business-domain observation, or corroborating DNS facts), retain the algorithm/version and uncertainty reasons, and remain suggestion-only. Candidate generation must reject cross-tenant records, public-suffix-only values, malformed or IP-only inputs, and suppressed/merged targets as applicable; it must not auto-attach a domain or infer ownership from a shared, parked, wildcard, sibling-subdomain, homograph, or merely resolvable domain. Every candidate needs human review, provenance, and an auditable accept/reject decision.
The platform must not claim that a domain is available, unregistered, or safe to acquire without an explicitly authorized availability provider registered with current product/legal/security approval, terms, tenant scope, rate limits, retention, and operational enablement. DNS nxdomain or no_data is not an availability result. Provider outages, rate limits, stale responses, conflicting results, and unknown status must remain unknown/unavailable, fail closed, and never trigger purchase, outreach, or automated follow-up.
Phase 7 remains a documentation/contract boundary in this MVP: there is no live DNS resolver, PSL-backed enrichment worker, cache service, or availability provider in Compose. Production work still includes selecting and versioning the PSL, implementing bounded DNS resolution and TTL-aware cache invalidation, defining MX/NS/TXT parsing and uncertainty retention, adding association review/permission/audit tests, and completing an approved availability-provider integration with SSRF/network egress controls, monitoring, retention, and incident/rollback procedures.
Phase 8 website scanning boundary
Website scanning is a bounded, tenant-scoped observation—not a crawler, browser, verifier, or outreach mechanism. A scan may fetch only http and https URLs after strict parsing and normalization. It must reject credentials, non-web schemes (file:, ftp:, gopher:, data:, javascript:, and similar), malformed hosts, localhost, IP literals where policy disallows them, and targets in loopback, private, link-local, multicast, reserved, or cloud-metadata ranges. DNS is resolved immediately before connection and the destination is revalidated at connection time; every redirect is limited, normalized, and revalidated for protocol, hostname, DNS, and IP range before it is followed. DNS answers must not be trusted from the initial validation alone (including rebinding changes).
Each scan enforces hard budgets: total wall-clock/request time, response bytes, body bytes retained, redirect count, and page/link crawl count and depth. Budgets apply across redirects and discovered links, with bounded concurrency, retries, and response decompression; a limit, timeout, DNS error, unsupported content type, or partial fetch produces an explicit incomplete/unknown outcome rather than an empty result. The scanner fetches HTML and other explicitly allowed small resources only; it does not submit forms, send credentials, execute JavaScript, load browser plugins, or perform arbitrary subresource requests.
Classifications are conservative and explainable. unknown, blocked, timeout, partial, and error remain distinct from a positive observation. A page can be classified only from bounded fetched content and must retain URL, redirect chain, response metadata, observed time, scanner/policy version, limits, and uncertainty reasons. A detected contact form, script, tracking tag, or business phrase is an observation—not proof of ownership, consent, deliverability, safety, or permission to contact.
Scan history is tenant-scoped and append-oriented. Results and cache entries are keyed by normalized URL plus scanner/policy/version inputs, bounded by size and retention, and expose observed_at, freshness/expiry, and whether a result came from cache. A cache hit is never represented as a fresh scan; policy, DNS, or scanner-version changes require revalidation/invalidation. History must not leak response bodies, secrets, cookies, authorization headers, or unnecessary personal data across tenants.
The website scanner remains a pilot boundary. Compose does not provide a production egress proxy, durable scan queue, distributed crawl coordinator, hardened DNS resolver, or compliance-grade result store. Production still requires independent SSRF testing (including DNS rebinding and redirect chains), egress/network policy, resource isolation, durable retention/deletion, authenticated scan-history authorization, rate limits and abuse controls, observability, and a reviewed policy for content types, robots/terms, caching, and incident response. Scans must never trigger acquisition, verification, enrichment, or outreach automatically.
Phase 9 public official-site contact extraction boundary
Phase 9 adds a passive, suggestion-only contact-observation workflow. When explicitly enabled, extraction may inspect bounded HTML from the business's approved/public official-site origin and its same-site contact/about pages; it is not general web search, crawling, enrichment, identity verification, or outreach. Only public page content and explicitly permitted mailto:/visible contact values may be considered. Do not submit forms, authenticate, bypass access controls, probe SMTP, send test messages, or contact a person or organization.
Every candidate contact must retain provenance: source URL and page location/context, extraction method, observed time, scanner/extractor and policy versions, and the exact uncertainty/reason code. Confidence is an explainable review signal, not deliverability, consent, ownership, or permission to contact. Classify role addresses separately from person addresses and classify free-mail domains separately from business-domain addresses; neither classification is proof of identity. Syntax validation is only a parse result. MX/DNS status is independently uncertain (not_checked, resolved, nxdomain, no_data, timeout, servfail, blocked, or error), and no MX result may be presented as deliverability.
False-positive exclusions must reject or quarantine values from asset URLs, image/file names, scripts/styles, example/test/placeholder domains, documentation text, tracking addresses, and malformed or unsupported schemes. Apply tenant-scoped suppressions before a candidate is persisted, returned, exported, or queued for review; suppressed values remain do-not-contact and suppression always wins over confidence, role, syntax, MX, pipeline, or verification state. Extraction is bounded by per-request and aggregate page/URL, byte, time, redirect, candidate, and concurrency limits. Store only the minimum contact value and lineage required for review, apply a documented retention/deletion class, and redact secrets and unnecessary personal data from logs and audit events.
Phase 9 does not authorize automated outreach. There is no SMTP probing, SMTP banner/VRFY/EXPN check, email validation message, send endpoint, campaign queue, or follow-up action. An extracted address is an observation requiring human review and explicit policy authorization before any separate future contact workflow.
Phase 10 configurable scoring boundary
Phase 10 separates fit score, priority band, and contact eligibility. A score is a deterministic ranking signal; it never authorizes contact. Rules are represented by a named, versioned rule set with explicit weights, thresholds, band definitions, eligibility gates, freshness policy, and suppression behavior. The active rule-set identifier/version is stored with each result so a historical score can be explained without silently applying today's policy.
A reproducible calculation uses the tenant-scoped business snapshot, normalized values, eligible evidence observations, rule-set/version, algorithm/version, and calculation time/freshness inputs. Explanations must retain the contributing factors, normalized inputs or evidence references, weights/points, exclusions, uncertainty reasons, and the final band. Do not accept a client-submitted score, band, eligibility flag, or rule version as authoritative.
Priority bands are policy labels (for example, high/medium/low or an explicitly configured equivalent) and must be derived from the versioned thresholds. Eligibility is evaluated separately and fail-closed: suppression/do-not-contact, stale or expired required evidence, unresolved/uncertain required signals, missing policy prerequisites, and authorization/tenant failures can make a prospect ineligible regardless of score. suppressed always wins and must remain visible; stale and uncertain observations must not be silently treated as absent or positive.
Recalculation is an explicit, tenant-scoped operation. It must snapshot the input/rule versions, record before/after score, band, eligibility, explanation, actor/job, timestamp, and reason in the audit trail, and be idempotent or safely repeatable. A policy/rule change must not rewrite history without an auditable recalculation; partial or failed recalculation must report its incomplete state rather than presenting mixed results as current.
Phase 10 remains a pilot boundary unless the runtime exposes all of the above controls end to end. Production work includes administrative rule-set lifecycle/approval, immutable calculation inputs, deterministic rounding/tie-breaking, scheduled recalculation with leases, retention and export semantics for explanations/audit, and regression tests proving suppression, stale, uncertain, and cross-tenant isolation behavior. See the API, security, and operations contracts for the authoritative safeguards.
Verification
python3 -m unittest discover -v -s apps/api/tests -t apps/api
python3 -m compileall -q apps/api apps/web
git diff --check
docker compose config --quiet
See apps/api/README.md, apps/web/README.md, docs/SECURITY.md, and docs/OPERATIONS.md for details.