- Type
- ${esc(type)}
- Email class
- ${esc(emailClass)}
- Validation
- ${esc(validation)}
- Suppression
- ${esc(suppression)}
- Source URL
- ${esc(source)}
diff --git a/README.md b/README.md index d249a73..896f334 100644 --- a/README.md +++ b/README.md @@ -120,6 +120,16 @@ Scan history is tenant-scoped and append-oriented. Results and cache entries are The website scanner remains a pilot boundary. Compose does not provide a production egress proxy, durable scan queue, distributed crawl coordinator, hardened DNS resolver, or compliance-grade result store. Production still requires independent SSRF testing (including DNS rebinding and redirect chains), egress/network policy, resource isolation, durable retention/deletion, authenticated scan-history authorization, rate limits and abuse controls, observability, and a reviewed policy for content types, robots/terms, caching, and incident response. Scans must never trigger acquisition, verification, enrichment, or outreach automatically. +## Phase 9 public official-site contact extraction boundary + +Phase 9 adds a **passive, suggestion-only** contact-observation workflow. When explicitly enabled, extraction may inspect bounded HTML from the business's approved/public official-site origin and its same-site contact/about pages; it is not general web search, crawling, enrichment, identity verification, or outreach. Only public page content and explicitly permitted `mailto:`/visible contact values may be considered. Do not submit forms, authenticate, bypass access controls, probe SMTP, send test messages, or contact a person or organization. + +Every candidate contact must retain provenance: source URL and page location/context, extraction method, observed time, scanner/extractor and policy versions, and the exact uncertainty/reason code. Confidence is an explainable review signal, not deliverability, consent, ownership, or permission to contact. Classify role addresses separately from person addresses and classify free-mail domains separately from business-domain addresses; neither classification is proof of identity. Syntax validation is only a parse result. MX/DNS status is independently uncertain (`not_checked`, `resolved`, `nxdomain`, `no_data`, `timeout`, `servfail`, `blocked`, or `error`), and no MX result may be presented as deliverability. + +False-positive exclusions must reject or quarantine values from asset URLs, image/file names, scripts/styles, example/test/placeholder domains, documentation text, tracking addresses, and malformed or unsupported schemes. Apply tenant-scoped suppressions before a candidate is persisted, returned, exported, or queued for review; suppressed values remain do-not-contact and suppression always wins over confidence, role, syntax, MX, pipeline, or verification state. Extraction is bounded by per-request and aggregate page/URL, byte, time, redirect, candidate, and concurrency limits. Store only the minimum contact value and lineage required for review, apply a documented retention/deletion class, and redact secrets and unnecessary personal data from logs and audit events. + +Phase 9 does not authorize automated outreach. There is no SMTP probing, SMTP banner/VRFY/EXPN check, email validation message, send endpoint, campaign queue, or follow-up action. An extracted address is an observation requiring human review and explicit policy authorization before any separate future contact workflow. + ## Verification ```bash diff --git a/apps/api/README.md b/apps/api/README.md index 6369ae1..e2469e2 100644 --- a/apps/api/README.md +++ b/apps/api/README.md @@ -1,4 +1,4 @@ -# Prospect Platform API — Phase 8 boundary +# Prospect Platform API — Phase 9 boundary Dependency-light JSON API for tenant-scoped prospect workflows and the Phase 8 bounded website-scanning, Phase 7 domain-intelligence, Phase 6 normalization/deduplication, and Phase 5 source-ingestion contracts. Core domain rules use Python's standard library and persistence is SQLite. Scan requests/results, where enabled, must remain auditable and fail closed; scanning never submits forms, executes JavaScript, or authorizes outreach. @@ -68,6 +68,8 @@ Implementations should expose source/query/job state without leaking raw payload The business detail includes the supported child collections: `contacts`, `domains`, `websites`, `evidence`, `pipeline`, and `notes`. The child collection routes are: - `POST /api/v1/businesses/{id}/contacts` — manually add a contact; suppression matching marks a matching contact as do-not-contact. +- `GET /api/v1/businesses/{id}/contacts/extract` — read the tenant-scoped Phase 9 official-site extraction projection. +- `POST /api/v1/businesses/{id}/contacts/extract` — extract bounded public contacts from the approved official site/same-site pages and persist provenance-bearing observations; the request is passive and must not probe SMTP or send outreach. Suppression matching marks matches `suppressed`/`do_not_contact` and remains authoritative. - `POST /api/v1/businesses/{id}/domains` — manually add a domain observation. - `POST /api/v1/businesses/{id}/websites` — manually add a website observation/classification. - `POST /api/v1/businesses/{id}/evidence` — manually add evidence with its kind, claim, and source URL/reference. This records provenance supplied by the operator; it does not scan or independently verify the URL. @@ -126,6 +128,16 @@ Classifications must be conservative, explainable, and derived only from bounded Scan history and cache reads/writes require the same tenant predicate as business routes. Keys include normalized URL, scanner/policy version, and relevant request/redirect policy; entries are size- and retention-bounded, expose `observed_at` and freshness/expiry, and never make a cache hit look like a fresh scan. Invalidate or re-evaluate entries after policy, DNS, or scanner-version changes. No scan result may trigger enrichment, acquisition, verification, or outreach. +## Phase 9 official-site contact extraction contract + +The optional Phase 9 extractor is a passive, authenticated, tenant-scoped observation of a business's approved/public official-site origin. It may inspect bounded HTML and same-site contact/about pages only; it must not become a search engine, unrestricted crawler, or arbitrary URL fetcher. The extractor must use the existing SSRF-safe URL, redirect, content-type, and resource-budget controls, and must never submit forms, execute target JavaScript, use credentials/cookies, probe SMTP, issue SMTP `VRFY`/`EXPN`, send validation mail, or perform outreach. + +For each candidate, return/store the normalized address only with provenance (official-site/page URL, page or DOM context, extraction method, observed time, extractor/policy version) and an explainable confidence/reason list. Preserve uncertainty rather than inventing facts. `syntax_valid`/`syntax_invalid` is a parser outcome only. Role classification (`role`/`person`/`unknown`) and free-mail classification (`free_mail`/`business_domain`/`unknown`) are independent review labels; they do not prove identity, consent, ownership, or deliverability. MX/DNS is a separate observation with resolver/source, observed time, TTL/freshness where available, and one of `not_checked`, `resolved`, `nxdomain`, `no_data`, `timeout`, `servfail`, `blocked`, or `error`; MX absence or failure remains unknown and must never be treated as invalid or undeliverable. + +Exclude false positives before persistence and response: values in scripts/styles/comments or asset URLs/file names, example/test/placeholder domains, tracking/telemetry addresses, malformed schemes, and unrelated third-party pages. Enforce hard limits for total extraction time, pages/URLs, redirects, response and retained bytes, candidates per page/request, and concurrency. Suppression matching is server-side and tenant-scoped, before storing, returning, exporting, or presenting a candidate; suppressed contacts are marked do-not-contact and cannot be revived by a later classification or review action. Retention must be explicit and bounded for extracted values, page provenance, DNS/MX observations, caches, and audit records; logs must not contain full contact payloads when a redacted identifier is sufficient. + +Extraction results are suggestions only and do not create a send/contact capability. The API exposes no SMTP-probe, validation-message, outreach, or campaign endpoint. If the feature is disabled, unapproved, over limit, blocked, or uncertain, fail closed with an explicit status/reason rather than an empty successful result. The current MVP remains pilot-only until extraction limits, suppression enforcement, retention/deletion jobs, provenance/audit coverage, and tenant-isolation tests are production hardened. + ## Remaining limitations and production migration work SQLite is a pilot store with an MVP job/event schema but no production migration runner, durable queue, scheduler, worker lease/recovery, durable backup command, or tested restore workflow. Authentication currently uses a development password fallback and does not by itself provide production Argon2id, MFA, CSRF protection, rate limiting, or a complete retention-grade audit system. Website scanning remains subject to pilot limits: no production egress proxy/isolation, distributed crawl coordinator, hardened resolver, or compliance-grade scan-history retention. Production work must add SSRF/DNS-rebinding/redirect-chain tests, egress policy, authenticated history/cache isolation, budget/abuse enforcement, durable result retention/deletion, observability, reviewed content/robots/terms policy, and the Phase 7 PSL/DNS/availability controls. Redis/Celery/Postgres remain future options, not implemented dependencies. diff --git a/apps/api/app/contact_extractor.py b/apps/api/app/contact_extractor.py new file mode 100644 index 0000000..c9258e9 --- /dev/null +++ b/apps/api/app/contact_extractor.py @@ -0,0 +1,139 @@ +"""Conservative extraction of public business contact signals from approved HTML. + +This module never fetches URLs. Callers must provide HTML obtained from the approved +website scanner and a source URL already associated with the business. +""" +from __future__ import annotations + +import html as html_lib +import re +from html.parser import HTMLParser +from urllib.parse import unquote, urlparse + +MAX_HTML_BYTES = 512 * 1024 +MAX_RESULTS = 100 +EMAIL_RE = re.compile(r"(? str: + return re.sub(r"\s+", " ", html_lib.unescape(value or "")).strip() + + +def _decode_obfuscation(value: str) -> str: + value = html_lib.unescape(unquote(value or "")) + value = re.sub(r"\s*(?:\[|\(|\{|\s)at(?:\]|\)|\}|\s)\s*", "@", value, flags=re.I) + value = re.sub(r"\s*(?:\[|\(|\{|\s)dot(?:\]|\)|\}|\s)*", ".", value, flags=re.I) + return value + + +def valid_email(value: str) -> bool: + value = value.strip().lower() + if len(value) > 254 or value.count("@") != 1 or value.split("@", 1)[1] in EXAMPLE_DOMAINS: + return False + return bool(EMAIL_RE.fullmatch(value)) and not any(x in value for x in ("password", "token", "secret", "credential", "apikey")) + + +def normalize_phone(value: str) -> str: + value = value.strip() + digits = re.sub(r"\D", "", value) + if value.startswith("+") and digits: + return "+" + digits + return digits + + +class _ContactParser(HTMLParser): + def __init__(self): + super().__init__(convert_charrefs=True) + self.text: list[str] = [] + self.links: list[tuple[str, str]] = [] + self.form_fields: list[str] = [] + self._ignore = 0 + self._anchor = "" + self._in_form = False + + def handle_starttag(self, tag, attrs): + attrs = {str(k).lower(): str(v or "") for k, v in attrs} + tag = tag.lower() + if tag in {"script", "style", "noscript", "template", "svg"}: + self._ignore += 1 + if tag == "a": self._anchor = attrs.get("href", "") + if tag == "form": self._in_form = True + if self._in_form and tag in {"input", "textarea", "select"}: + name = attrs.get("name", "") or attrs.get("id", "") + typ = attrs.get("type", "") + self.form_fields.append(_clean(" ".join((name, typ, attrs.get("placeholder", ""))))) + + def handle_endtag(self, tag): + tag = tag.lower() + if tag in {"script", "style", "noscript", "template", "svg"} and self._ignore: self._ignore -= 1 + if tag == "a": self._anchor = "" + if tag == "form": self._in_form = False + + def handle_data(self, data): + if self._ignore: return + if data.strip(): self.text.append(data) + if self._anchor: self.links.append((self._anchor, data)) + + +def _record(kind, value, label, source_url, confidence, classification="unknown", *, suppressed=False, provenance="visible_text"): + return {"kind": kind, "value": value, "label": label[:200], "classification": classification, + "confidence": round(max(0.0, min(1.0, confidence)), 2), "source_url": source_url, + "public_business": True, "mx_status": "unknown", "suppressed": bool(suppressed), + "do_not_contact": bool(suppressed), "provenance": provenance} + + +def extract_contacts(source_html: str, source_url: str, *, suppressions=None, max_results=MAX_RESULTS) -> list[dict]: + if not isinstance(source_html, str) or len(source_html.encode("utf-8")) > MAX_HTML_BYTES: + raise ValueError("html_too_large") + try: max_results = int(max_results) + except (ValueError, TypeError): raise ValueError("invalid_limits") + if max_results < 1 or max_results > MAX_RESULTS: raise ValueError("invalid_limits") + parser = _ContactParser(); parser.feed(source_html) + visible = _decode_obfuscation(_clean(" ".join(parser.text))) + suppression = {(str(x.get("kind", "")), str(x.get("value", "")).lower()) for x in (suppressions or [])} + out, seen = [], set() + def add(kind, value, label, confidence, classification="unknown", provenance="visible_text"): + value = value.strip().lower() if kind == "email" else value.strip() + if kind == "email": + if not valid_email(value): return + local, domain = value.rsplit("@", 1) + classification = "free_mail" if domain in FREE_MAIL else ("role" if local in ROLE_NAMES else "named") + key = (kind, value); blocked = ("email", value) in suppression or ("domain", domain) in suppression + else: + if kind in {"phone", "whatsapp"}: value = normalize_phone(value) + if len(re.sub(r"\D", "", value)) < 7: return + key = (kind, value); blocked = (kind, value.lower()) in suppression + if key in seen or len(out) >= max_results: return + seen.add(key); out.append(_record(kind, value, label or kind.title(), source_url, confidence, classification, suppressed=blocked, provenance=provenance)) + for href, anchor_text in parser.links: + raw = _decode_obfuscation(href) + if raw.lower().startswith("mailto:"): + address = raw[7:].split("?", 1)[0] + add("email", address, _clean(anchor_text), 0.98, provenance="mailto") + elif raw.lower().startswith("tel:"): + add("phone", raw[4:].split("?", 1)[0], _clean(anchor_text), 0.98, provenance="tel") + else: + parsed = urlparse(raw if "://" in raw else "https://" + raw) + host = (parsed.hostname or "").lower().removeprefix("www.") + if host == "wa.me" or "whatsapp" in host: + number = re.sub(r"\D", "", parsed.path) + if number: add("whatsapp", "+" + number, _clean(anchor_text) or "WhatsApp", 0.98, provenance="whatsapp_link") + elif any(host == d or host.endswith("." + d) for d in SOCIAL_HOSTS): + add("social", raw, _clean(anchor_text) or host, 0.95, provenance="social_link") + for match in EMAIL_RE.finditer(visible): + context = visible[max(0, match.start() - 40):match.start()] + if re.search(r"(?:password|passwd|token|secret|credential|api[_ -]?key|authorization)\s*[:=]?\s*$", context, re.I): + continue + add("email", match.group(0), "Email", 0.88) + for value in PHONE_RE.findall(visible): add("phone", value, "Phone", 0.82) + # A form is provenance, not proof of a destination address. + if parser.form_fields: + for field in parser.form_fields[:5]: + if re.search(r"email|contact|phone|whatsapp", field, re.I): + add("form", field, "Contact form", 0.7, provenance="form_field") + return out diff --git a/apps/api/app/main.py b/apps/api/app/main.py index 863c49c..3e4bd60 100644 --- a/apps/api/app/main.py +++ b/apps/api/app/main.py @@ -11,11 +11,13 @@ if __package__ in (None, ""): from app.sources import adapter_for, contains_secret from app.domain_intelligence import normalize_registrable_domain, resolve_domain, generate_candidate_domains from app.website_scanner import scan_website, validate_url + from app.contact_extractor import extract_contacts, MAX_HTML_BYTES, MAX_RESULTS else: from .domain import deduplication_key, deduplicate_businesses, is_suppressed, normalize_business, score_business, normalize_domain, normalize_phone, match_businesses from .sources import adapter_for, contains_secret from .domain_intelligence import normalize_registrable_domain, resolve_domain, generate_candidate_domains from .website_scanner import scan_website, validate_url + from .contact_extractor import extract_contacts, MAX_HTML_BYTES, MAX_RESULTS ORGANIZATION_ID = "demo-tenant" SCHEMA = Path(__file__).resolve().parents[1] / "schema.sql" SESSION_DAYS = 7 @@ -25,6 +27,7 @@ JOB_TYPES = {"noop", "prospect_recalculate", "source_discovery", "domain_check"} JOB_PAGE_SIZE = 100 WEBSITE_SCAN_PAGE_SIZE = 100 WEBSITE_SCAN_CACHE_SECONDS = 3600 +CONTACT_EXTRACTION_PAGE_SIZE = 100 SECRET_KEYS = {"password", "passwd", "secret", "token", "api_key", "apikey", "authorization", "credential", "private_key"} CHILD_TABLES = {"contacts": ("name", "email", "phone", "title", "do_not_contact"), "domains": ("domain", "kind"), "websites": ("url", "website_class"), "evidence": ("kind", "url", "claim"), "notes": ("body",)} @@ -125,6 +128,7 @@ class ApiHandler(BaseHTTPRequestHandler): if path=="/api/v1/jobs": return self.list_jobs(db,org,parse_qs(parsed.query)) if path=="/api/v1/domain-checks": return self.list_domain_checks(db,org,parse_qs(parsed.query)) if path=="/api/v1/website-scans": return self.list_website_scans(db,org,parse_qs(parsed.query)) + if path=="/api/v1/contact-extractions": return self.list_contact_extractions(db,org,parse_qs(parsed.query)) if path.startswith("/api/v1/jobs/"): return self.get_job_route(db,org,path,parse_qs(parsed.query)) if path.startswith("/api/v1/businesses/"): bits=path.split("/"); ident=bits[4] if len(bits)>4 else "" @@ -133,6 +137,7 @@ class ApiHandler(BaseHTTPRequestHandler): if not row:return self.send_json(404,{"error":"not_found"}) if len(bits)==7 and bits[5:]==["websites","scan"]: return self.get_latest_website_scan(int(ident),db,user) if len(bits)==7 and bits[5:]==["domains","check"]: return self.get_domain_check(int(ident),db,user,parse_qs(parsed.query)) + if len(bits)==7 and bits[5:]==["contacts","extract"]: return self.send_json(405,{"error":"method_not_allowed"}) if len(bits)==6 and bits[5]=="domain-candidates": return self.list_domain_candidates(int(ident),db,org) if len(bits)==6 and bits[5]=="matches": return self.matches(int(ident),db,org) payload=row_json(row); payload.update(self.nested(db,int(ident),org)); return self.send_json(200,payload) @@ -194,6 +199,64 @@ class ApiHandler(BaseHTTPRequestHandler): rows = db.execute("SELECT * FROM website_scans WHERE " + " AND ".join(where) + " ORDER BY id DESC LIMIT ? OFFSET ?", params + [limit + 1, offset]).fetchall() return self.send_json(200, {"organization_id": org, "items": [self._website_scan_result(r) for r in rows[:limit]], "limit": limit, "offset": offset, "has_more": len(rows) > limit}) + def _contact_json(self, row): + result = row_json(row) + for key in ("public_business", "suppressed", "do_not_contact"): + if key in result: result[key] = bool(result[key]) + return result + + def list_contact_extractions(self, db, org, query): + try: + limit = int((query.get("page_size") or [50])[0]); offset = max(0, int((query.get("offset") or [0])[0])) + if limit < 1 or limit > CONTACT_EXTRACTION_PAGE_SIZE: raise ValueError + except (ValueError, TypeError): return self.send_json(400, {"error": "invalid_pagination"}) + params = [org]; where = ["organization_id=?"] + business_id = (query.get("business_id") or [""])[0] + if business_id.isdigit(): where.append("business_id=?"); params.append(int(business_id)) + rows = db.execute("SELECT * FROM contact_extractions WHERE " + " AND ".join(where) + " ORDER BY id DESC LIMIT ? OFFSET ?", params + [limit + 1, offset]).fetchall() + return self.send_json(200, {"organization_id": org, "items": [self._contact_json(r) for r in rows[:limit]], "limit": limit, "offset": offset, "has_more": len(rows) > limit}) + + def extract_business_contacts(self, bid, payload, db, user): + org = user["organization_id"]; business = self.business(db, bid, org) + if not business: return self.send_json(404, {"error": "not_found"}) + scan_id = payload.get("website_scan_id", payload.get("scan_id")) + scan = None + if scan_id is not None: + if not isinstance(scan_id, int): return self.send_json(400, {"error": "invalid_scan"}) + scan = db.execute("SELECT * FROM website_scans WHERE id=? AND business_id=? AND organization_id=?", (scan_id, bid, org)).fetchone() + if not scan: return self.send_json(404, {"error": "scan_not_found"}) + try: stored = json.loads(scan["result_json"] or "{}") + except (TypeError, ValueError): stored = {} + source_url = str(payload.get("source_url") or scan["input_url"]).strip() + source_html = payload.get("html") if isinstance(payload.get("html"), str) else stored.get("html") + if source_html is None: return self.send_json(409, {"error": "scan_html_unavailable"}) + if source_url != scan["input_url"] and source_url != (stored.get("final_url") or ""): return self.send_json(400, {"error": "source_not_approved"}) + else: + source_url = str(payload.get("source_url") or "").strip() + source_html = payload.get("html") + if not source_url or not isinstance(source_html, str): return self.send_json(400, {"error": "approved_scan_required"}) + parsed = urlparse(source_url) + official_hosts = {str(business["website_domain"]).lower().strip(".")} + if business["website"]: + official_hosts.add((urlparse(business["website"]).hostname or "").lower().strip(".")) + if parsed.scheme not in {"http", "https"} or not parsed.hostname or parsed.hostname.lower().strip(".") not in official_hosts: + return self.send_json(400, {"error": "source_not_approved"}) + if len(source_html.encode("utf-8")) > MAX_HTML_BYTES: return self.send_json(413, {"error": "html_too_large"}) + try: requested_limit = int(payload.get("limit", MAX_RESULTS)) + except (ValueError, TypeError): return self.send_json(400, {"error": "invalid_limits"}) + if requested_limit < 1 or requested_limit > MAX_RESULTS: return self.send_json(400, {"error": "invalid_limits"}) + key = str(payload.get("idempotency_key") or hashlib.sha256((str(scan_id or "") + source_url + source_html).encode()).hexdigest())[:200] + existing = db.execute("SELECT * FROM contact_extractions WHERE organization_id=? AND business_id=? AND extraction_key=? ORDER BY id", (org, bid, key)).fetchall() + if existing: return self.send_json(200, {"business_id": bid, "extraction_key": key, "items": [self._contact_json(r) for r in existing], "idempotent": True}) + suppressions = [dict(r) for r in db.execute("SELECT kind,value FROM suppressions WHERE organization_id=?", (org,))] + try: found = extract_contacts(source_html, source_url, suppressions=suppressions, max_results=requested_limit) + except ValueError as exc: return self.send_json(400, {"error": str(exc)}) + for item in found: + db.execute("INSERT INTO contact_extractions(organization_id,business_id,website_scan_id,extraction_key,kind,value,label,classification,confidence,source_url,public_business,mx_status,suppressed,do_not_contact,provenance) VALUES(?,?,?,?,?,?,?,?,?,?,?,?,?,?,?)", (org,bid,scan["id"] if scan else None,key,item["kind"],item["value"],item["label"],item["classification"],item["confidence"],item["source_url"],int(item["public_business"]),item["mx_status"],int(item["suppressed"]),int(item["do_not_contact"]),item["provenance"])) + self.audit(db, user, "contacts.extracted", f"{bid}:{len(found)}:{key}"); db.commit() + rows = db.execute("SELECT * FROM contact_extractions WHERE organization_id=? AND business_id=? AND extraction_key=? ORDER BY id", (org,bid,key)).fetchall() + return self.send_json(201, {"business_id": bid, "extraction_key": key, "items": [self._contact_json(r) for r in rows], "idempotent": False}) + def get_latest_website_scan(self, bid, db, user): row = db.execute("SELECT * FROM website_scans WHERE business_id=? AND organization_id=? ORDER BY id DESC LIMIT 1", (bid, user["organization_id"])).fetchone() if not row: return self.send_json(404, {"error": "scan_not_found"}) @@ -333,8 +396,10 @@ class ApiHandler(BaseHTTPRequestHandler): if path=="/api/v1/sources":return self.create_source(payload,db,user) if path=="/api/v1/discovery-queries":return self.create_query(payload,db,user) if path=="/api/v1/suppressions":return self.create_suppression(payload,db,user) + if path=="/api/v1/contact-extractions": return self.send_json(405,{"error":"method_not_allowed"}) if path=="/api/v1/imports/preview":return self.preview_import(payload,db,org) if len(path.split("/"))==7 and path.split("/")[3:6]==["businesses",path.split("/")[4],"websites"] and path.split("/")[6]=="scan": return self.scan_business_website(int(path.split("/")[4]) if path.split("/")[4].isdigit() else -1,payload,db,user) + if len(path.split("/"))==7 and path.split("/")[3:6]==["businesses",path.split("/")[4],"contacts"] and path.split("/")[6]=="extract": return self.extract_business_contacts(int(path.split("/")[4]) if path.split("/")[4].isdigit() else -1,payload,db,user) if len(path.split("/"))==7 and path.split("/")[3:6]==["businesses",path.split("/")[4],"domains"] and path.split("/")[6]=="check": return self.post_domain_check(int(path.split("/")[4]) if path.split("/")[4].isdigit() else -1,payload,db,user) if len(path.split("/"))==7 and path.split("/")[3:6]==["businesses",path.split("/")[4],"domain-candidates"] and path.split("/")[6]=="check-availability": return self.check_availability(int(path.split("/")[4]) if path.split("/")[4].isdigit() else -1,payload,db,user) if path.startswith("/api/v1/merge-history/") and path.endswith("/reverse"): diff --git a/apps/api/app/website_scanner.py b/apps/api/app/website_scanner.py index ce0f824..db0da02 100644 --- a/apps/api/app/website_scanner.py +++ b/apps/api/app/website_scanner.py @@ -169,6 +169,8 @@ def scan_website(url: str, *, timeout: float = DEFAULT_TIMEOUT, max_bytes: int = if not 0 < int(max_redirects) <= MAX_REDIRECTS or not 0 < int(max_pages) <= MAX_PAGES: raise ValueError("invalid_limits") fetched = _fetch(url, timeout=max(0.1, min(float(timeout), 10.0)), max_bytes=max(1, min(int(max_bytes), MAX_BYTES)), max_redirects=int(max_redirects)) result.update({k: fetched[k] for k in ("status", "final_url", "redirect_chain", "elapsed_ms", "tls", "certificate_status")}); result["size_bytes"] = len(fetched["body"]) + if fetched["content_type"] in {"text/html", "application/xhtml+xml"}: + result["html"] = fetched["body"].decode("utf-8", "replace") if fetched["content_type"] not in {"text/html", "application/xhtml+xml"}: result["classification"] = classify_website(fetched["status"], fetched["final_url"], ""); return result parser = _PageParser(); parser.feed(fetched["body"].decode("utf-8", "replace")) diff --git a/apps/api/schema.sql b/apps/api/schema.sql index fca636b..fd42d32 100644 --- a/apps/api/schema.sql +++ b/apps/api/schema.sql @@ -188,3 +188,27 @@ CREATE TABLE IF NOT EXISTS website_scans ( CREATE INDEX IF NOT EXISTS idx_website_scans_org ON website_scans(organization_id,created_at DESC,id DESC); CREATE INDEX IF NOT EXISTS idx_website_scans_business ON website_scans(organization_id,business_id,created_at DESC,id DESC); CREATE INDEX IF NOT EXISTS idx_website_scans_cache ON website_scans(organization_id,cache_key,cache_expires_at); + +-- Phase 9 contact extraction provenance. MX is intentionally a state value, not a probe. +CREATE TABLE IF NOT EXISTS contact_extractions ( + id INTEGER PRIMARY KEY AUTOINCREMENT, + organization_id TEXT NOT NULL REFERENCES organizations(id), + business_id INTEGER NOT NULL REFERENCES businesses(id) ON DELETE CASCADE, + website_scan_id INTEGER REFERENCES website_scans(id) ON DELETE SET NULL, + extraction_key TEXT NOT NULL, + kind TEXT NOT NULL, + value TEXT NOT NULL, + label TEXT NOT NULL DEFAULT '', + classification TEXT NOT NULL DEFAULT 'unknown', + confidence REAL NOT NULL DEFAULT 0, + source_url TEXT NOT NULL, + public_business INTEGER NOT NULL DEFAULT 1, + mx_status TEXT NOT NULL DEFAULT 'unknown', + suppressed INTEGER NOT NULL DEFAULT 0, + do_not_contact INTEGER NOT NULL DEFAULT 0, + provenance TEXT NOT NULL DEFAULT '', + created_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP, + UNIQUE(organization_id,business_id,extraction_key,kind,value) +); +CREATE INDEX IF NOT EXISTS idx_contact_extractions_org ON contact_extractions(organization_id,created_at DESC,id DESC); +CREATE INDEX IF NOT EXISTS idx_contact_extractions_business ON contact_extractions(organization_id,business_id,id DESC); diff --git a/apps/api/tests/test_phase9_api.py b/apps/api/tests/test_phase9_api.py new file mode 100644 index 0000000..dd68fa6 --- /dev/null +++ b/apps/api/tests/test_phase9_api.py @@ -0,0 +1,52 @@ +import json +import os +import threading +import unittest +from http.client import HTTPConnection +from tempfile import TemporaryDirectory + +from app.main import create_server + + +class ContactExtractionApiTests(unittest.TestCase): + def setUp(self): + self.tmp = TemporaryDirectory() + os.environ['BOOTSTRAP_ADMIN_EMAIL'] = 'extract-owner@example.test' + os.environ['BOOTSTRAP_ADMIN_PASSWORD'] = 'password' + self.server = create_server('127.0.0.1', 0, self.tmp.name + '/db.sqlite') + self.thread = threading.Thread(target=self.server.serve_forever, daemon=True); self.thread.start() + self.conn = HTTPConnection('127.0.0.1', self.server.server_port, timeout=4); self.cookie = None + self.request('POST', '/api/v1/auth/login', {'email': 'extract-owner@example.test', 'password': 'password'}) + + def tearDown(self): + self.server.shutdown(); self.server.server_close(); self.thread.join(2); self.tmp.cleanup() + + def request(self, method, path, payload=None): + body = json.dumps(payload).encode() if payload is not None else None + headers = {'Content-Type': 'application/json'} if body else {} + if self.cookie: headers['Cookie'] = self.cookie + self.conn.request(method, path, body, headers); response = self.conn.getresponse() + cookie = response.getheader('Set-Cookie') + if cookie: self.cookie = cookie.split(';', 1)[0] + return response.status, json.loads(response.read() or b'{}') + + def test_official_html_extracts_with_provenance_suppression_and_idempotency(self): + _, business = self.request('POST', '/api/v1/businesses', {'name': 'Acme', 'website': 'https://acme.test'}) + self.request('POST', '/api/v1/suppressions', {'kind': 'email', 'value': 'sales@acme.test'}) + payload = {'source_url': 'https://acme.test/contact', 'html': 'Sales
info [at] acme [dot] test
', 'idempotency_key': 'extract-1'} + status, result = self.request('POST', f"/api/v1/businesses/{business['id']}/contacts/extract", payload) + self.assertEqual(status, 201); self.assertEqual(len(result['items']), 2) + sales = next(x for x in result['items'] if x['value'] == 'sales@acme.test') + self.assertTrue(sales['suppressed']); self.assertTrue(sales['do_not_contact']); self.assertEqual(sales['provenance'], 'mailto'); self.assertEqual(sales['mx_status'], 'unknown') + self.assertEqual(self.request('POST', f"/api/v1/businesses/{business['id']}/contacts/extract", payload)[1]['idempotent'], True) + self.assertEqual(self.request('GET', '/api/v1/contact-extractions?page_size=1')[1]['limit'], 1) + + def test_arbitrary_and_oversized_sources_are_rejected(self): + _, business = self.request('POST', '/api/v1/businesses', {'name': 'Acme', 'website': 'https://acme.test'}) + path = f"/api/v1/businesses/{business['id']}/contacts/extract" + self.assertEqual(self.request('POST', path, {'source_url': 'https://evil.test', 'html': 'x@y.test
'})[1]['error'], 'source_not_approved') + self.assertEqual(self.request('POST', path, {'source_url': 'https://acme.test', 'html': 'x' * (512 * 1024 + 1)})[0], 413) + + +if __name__ == '__main__': + unittest.main() diff --git a/apps/api/tests/test_phase9_extractor.py b/apps/api/tests/test_phase9_extractor.py new file mode 100644 index 0000000..fac2384 --- /dev/null +++ b/apps/api/tests/test_phase9_extractor.py @@ -0,0 +1,42 @@ +import unittest +from app.contact_extractor import extract_contacts + + +class ContactExtractorTests(unittest.TestCase): + def test_extracts_public_mailto_obfuscated_phone_and_ignores_false_positives(self): + html = ''' + Email Sales + info [at] acme.test+27 (12) 345-6789 + WhatsApp + +API key: foo@acme.test password: bar@acme.test
', 'https://acme.test/') + self.assertEqual(result, []) + + def test_classifies_role_named_free_mail_and_suppression(self): + html = 'support@acme.test alice@acme.test bob@gmail.com
' + result = extract_contacts(html, 'https://acme.test/', suppressions=[{'kind':'email','value':'support@acme.test'}]) + by = {x['value']: x for x in result} + self.assertEqual(by['support@acme.test']['classification'], 'role') + self.assertTrue(by['support@acme.test']['do_not_contact']) + self.assertTrue(by['support@acme.test']['suppressed']) + self.assertEqual(by['alice@acme.test']['classification'], 'named') + self.assertEqual(by['bob@gmail.com']['classification'], 'free_mail') + self.assertEqual(by['alice@acme.test']['mx_status'], 'unknown') + + +if __name__ == '__main__': + unittest.main() diff --git a/apps/web/README.md b/apps/web/README.md index 0eab151..6ae7ea4 100644 --- a/apps/web/README.md +++ b/apps/web/README.md @@ -1,4 +1,4 @@ -# ProspectOS web — Phase 8 boundary +# ProspectOS web — Phase 9 boundary Self-contained static frontend for the Prospect Platform API. There is no bundler or runtime dependency: serve this directory with any static HTTP server. The UI supports a manual, tenant-scoped review workflow, Phase 4 MVP job monitor, and Phase 8 scan-result/history presentation when supplied by the API. The browser does not fetch targets, submit forms, execute scan JavaScript, or send outreach; SSRF controls and budgets are server-side. @@ -66,6 +66,14 @@ The current client renders job status/counts, detail, structured errors, progres A zero-dependency static smoke page (`smoke-test.html`) checks the key DOM contract in an iframe and reports pass/fail. Browser smoke coverage should treat pagination, detail child records, provenance rendering, pipeline/notes actions, and the disabled outreach state as separate checks. +## Phase 9 contact-observation UI contract + +When the API supplies Phase 9 results, the UI may render contacts extracted from the approved/public official site and same-site contact/about pages. Display source URL/page context, extraction method, observed time, extractor/policy version, confidence and reasons, syntax status, role classification, free-mail classification, and independent MX/DNS status/freshness. Use explicit labels such as **Observation**, **Human review required**, and **Unknown**; never label a candidate verified, deliverable, owned, consented, or ready for outreach. The UI must show suppressed contacts as **Do not contact**, preserve the suppression reason, and never hide or override suppression through filters, refreshes, exports, or cached results. + +Extraction is not a browser crawler. The browser must not fetch target pages directly, submit contact forms, execute target JavaScript, send credentials/cookies, probe SMTP, send validation messages, or expose an outreach/send control. Render bounded/partial/blocked/timeout/error results distinctly from an empty successful result, including candidate/page/byte/time limit reasons. Do not display false-positive candidates from assets, scripts/styles, example/test/placeholder values, tracking addresses, malformed schemes, or unrelated third-party pages. + +The API remains authoritative for official-site scope, tenant isolation, suppression enforcement, limits, retention, provenance, and permissions. A cached extraction must show its observed time and freshness, never “live.” Candidate confidence, role/free-mail labels, syntax, and MX/DNS uncertainty are review metadata only and cannot enable a contact action. + ## Remaining limitations -The static client has no client-side crawler, scanner, enrichment scheduler, outreach integration, availability provider, or SSE delivery. It can display server-provided Phase 8 observations, but production still requires server-side SSRF/DNS-rebinding/redirect controls, hard size/time/crawl budgets, durable scan history/cache isolation and retention, abuse/rate controls, and authenticated permission/audit coverage. For domain intelligence, display registrable-domain/PSL version and unresolved reasons, DNS status and freshness/TTL (not “available”), independent MX/NS/TXT uncertainty, and association confidence with explainable evidence. Never auto-attach candidate domains or treat `nxdomain`/`no_data` as availability. CSV preview is capped for display and is not an import workflow. +The static client has no client-side crawler, scanner, contact extractor, enrichment scheduler, outreach integration, availability provider, or SSE delivery. It can display server-provided Phase 9 observations, but production still requires server-side official-site scoping, SSRF/DNS-rebinding/redirect controls, hard extraction/page/byte/time/candidate budgets, durable history/cache isolation and retention/deletion, abuse/rate controls, suppression regression tests, and authenticated provenance/audit coverage. For domain intelligence, display registrable-domain/PSL version and unresolved reasons, DNS status and freshness/TTL (not “available”), independent MX/NS/TXT uncertainty, and association confidence with explainable evidence. Never auto-attach candidate domains or treat `nxdomain`/`no_data` as availability. CSV preview is capped for display and is not an import workflow. diff --git a/apps/web/app.js b/apps/web/app.js index 90bdc7a..6ecd65e 100644 --- a/apps/web/app.js +++ b/apps/web/app.js @@ -74,7 +74,20 @@ return `${error?`${esc(error)}
${esc(error.message)}
PROSPECT DETAIL
${esc(p.website_domain||'no detected website')}
✓ ${esc(labelFactor(x))}${esc(p.confidence||'Medium')}
`).join(''):''}${esc(review)}
${st!=='suppressed'?'':''}${st==='suppressed'?'Suppressed records cannot be contacted.':'Review this prospect before outreach is available.'}
`:''}`;renderWebsiteScanPanel(p);renderDomainPanel(p);renderDedupPanel();} + const extractedContactItems = payload => { for (const key of ['contacts','extracted_contacts','items','results']) if (Array.isArray(payload?.[key])) return payload[key]; return Array.isArray(payload) ? payload : []; }; + const extractedContactValue = (item, keys, fallback='Unknown') => { for (const key of keys) if (item?.[key] !== undefined && item[key] !== null && item[key] !== '') return item[key]; return fallback; }; + const extractedContactStatus = item => { const status=String(extractedContactValue(item,['suppression_status','suppression','status'],'Unknown')).toLowerCase().replaceAll('_','-'); return status==='suppressed'||item?.suppressed===true ? 'Suppressed' : status==='not-suppressed'||status==='allowed'||item?.suppressed===false ? 'Not suppressed' : status.replaceAll('-',' ').replace(/\\b\\w/g,c=>c.toUpperCase()); }; + function renderExtractedContacts(payload){ + const state=$('contactExtractionState'); if(!state)return; + if(!payload){state.innerHTML='No extracted contacts returned. Extract from approved public business pages to review contact evidence.
';return;} + const contacts=extractedContactItems(payload); + if(!contacts.length){state.innerHTML='No contact details were found on the approved public business pages checked.
PUBLIC CONTACT EXTRACTION
Extraction uses approved public business pages only. It does not probe SMTP, verify mailbox access, or send outreach.
No public contacts extracted yet. Start an extraction to review evidence.
'}No public contacts extracted yet. Start an extraction to review evidence.
';return 'Previously extracted contacts are available. Refresh to check approved public business pages again.
';} + async function loadContactExtraction(id,{extract=false}={}){const state=$('contactExtractionState'),extractButton=$('extractContactsBtn'),refreshButton=$('refreshContactExtractionBtn');if(!state||!id)return;[extractButton,refreshButton].forEach(button=>{if(button)button.disabled=true;});state.innerHTML='${esc(error.message)}
PROSPECT DETAIL
${esc(p.website_domain||'no detected website')}
✓ ${esc(labelFactor(x))}${esc(p.confidence||'Medium')}
`).join(''):''}${esc(review)}
${st!=='suppressed'?'':''}${st==='suppressed'?'Suppressed records cannot be contacted.':'Review this prospect before outreach is available.'}
`:''}`;renderContactExtractionPanel(p);loadContactExtraction(p.id);renderWebsiteScanPanel(p);renderDomainPanel(p);renderDedupPanel();} let mergeSource = null, mergeTarget = null, mergeBusy = false; const payloadItems = (payload, keys) => { for (const key of keys) if (Array.isArray(payload?.[key])) return payload[key]; return Array.isArray(payload) ? payload : []; }; const suggestionId = item => item.target_id ?? item.business_id ?? item.prospect_id ?? item.id; @@ -161,7 +174,7 @@ async function logout(){try{await fetch(endpoint('/api/v1/auth/logout'),{method:'POST',credentials:'include'});}finally{showLogin('You have been signed out.');$('loginForm').reset();}} async function bootstrap(){try{const res=await fetch(endpoint('/api/v1/auth/me'),{credentials:'include'});if(res.status===401){showLogin();return;}if(!res.ok)throw new Error('Could not verify session.');const user=await res.json();showDashboard(user.user||user);await loadData();await loadJobs();await loadSources();}catch(e){if(e.message!=='unauthorized')showLogin('Unable to connect to the workspace. Try again.');}} document.addEventListener('submit',e=>{if(e.target.id==='contactForm')saveContact(e.target);if(e.target.id==='noteForm')saveNote(e.target);if(e.target.id==='pipelineForm')saveStage(e.target);}); - document.addEventListener('click',e=>{if(e.target.id==='verifyBtn')verify();if(e.target.id==='scanWebsiteBtn'&&selectedId)loadWebsiteScan(selectedId,{scan:true});if(e.target.id==='refreshWebsiteScanBtn'&&selectedId)loadWebsiteScan(selectedId);if(e.target.id==='retryWebsiteScanBtn'&&selectedId)loadWebsiteScan(selectedId);if(e.target.id==='runDomainCheckBtn')runDomainCheck();if(e.target.id==='retryDomainBtn'&&selectedId)loadDomainIntelligence(selectedId);const availability=e.target.closest?.('[data-domain-availability]');if(availability)checkDomainAvailability(availability.dataset.domain,availability);if(e.target.id==='retryDetailBtn'&&selectedId)loadDetail(selectedId);if(e.target.id==='retryDedupBtn'&&selectedId)loadMatchSuggestions(selectedId);if(e.target.id==='retryHistoryBtn'&&selectedId)loadMergeHistory(selectedId);if(e.target.id==='cancelMergeBtn'||e.target.id==='cancelMergeBtnSecondary')closeMergeDialog();if(e.target.id==='confirmMergeBtn')confirmMerge();const mergeButton=e.target.closest?.('[data-merge-target]');if(mergeButton)openMergeDialog(mergeButton.dataset.mergeTarget,mergeButton.dataset.mergeTargetName);const reverseButton=e.target.closest?.('[data-reverse-merge]');if(reverseButton)reverseMerge(reverseButton.dataset.reverseMerge);if(e.target.id==='retryJobDetailBtn'&&selectedJobId)loadJobDetail(selectedJobId);if(e.target.id==='cancelJobBtn')jobAction('cancel');if(e.target.id==='retryJobBtn')jobAction('retry');const row=e.target.closest?.('[data-job-id]');if(row)loadJobDetail(row.dataset.jobId);}); + document.addEventListener('click',e=>{if(e.target.id==='verifyBtn')verify();if(e.target.id==='extractContactsBtn'&&selectedId)loadContactExtraction(selectedId,{extract:true});if(e.target.id==='refreshContactExtractionBtn'&&selectedId)loadContactExtraction(selectedId,{extract:true});if(e.target.id==='retryContactExtractionBtn'&&selectedId)loadContactExtraction(selectedId);if(e.target.id==='scanWebsiteBtn'&&selectedId)loadWebsiteScan(selectedId,{scan:true});if(e.target.id==='refreshWebsiteScanBtn'&&selectedId)loadWebsiteScan(selectedId);if(e.target.id==='retryWebsiteScanBtn'&&selectedId)loadWebsiteScan(selectedId);if(e.target.id==='runDomainCheckBtn')runDomainCheck();if(e.target.id==='retryDomainBtn'&&selectedId)loadDomainIntelligence(selectedId);const availability=e.target.closest?.('[data-domain-availability]');if(availability)checkDomainAvailability(availability.dataset.domain,availability);if(e.target.id==='retryDetailBtn'&&selectedId)loadDetail(selectedId);if(e.target.id==='retryDedupBtn'&&selectedId)loadMatchSuggestions(selectedId);if(e.target.id==='retryHistoryBtn'&&selectedId)loadMergeHistory(selectedId);if(e.target.id==='cancelMergeBtn'||e.target.id==='cancelMergeBtnSecondary')closeMergeDialog();if(e.target.id==='confirmMergeBtn')confirmMerge();const mergeButton=e.target.closest?.('[data-merge-target]');if(mergeButton)openMergeDialog(mergeButton.dataset.mergeTarget,mergeButton.dataset.mergeTargetName);const reverseButton=e.target.closest?.('[data-reverse-merge]');if(reverseButton)reverseMerge(reverseButton.dataset.reverseMerge);if(e.target.id==='retryJobDetailBtn'&&selectedJobId)loadJobDetail(selectedJobId);if(e.target.id==='cancelJobBtn')jobAction('cancel');if(e.target.id==='retryJobBtn')jobAction('retry');const row=e.target.closest?.('[data-job-id]');if(row)loadJobDetail(row.dataset.jobId);}); $('loginForm').addEventListener('submit',login);$('logoutBtn').addEventListener('click',logout);$('searchInput').addEventListener('input',()=>{page=1;renderRows();});['scoreFilter','statusFilter','websiteClassFilter','pipelineFilter'].forEach(id=>$(id).addEventListener('change',()=>{page=1;loadData();}));$('pageSize').addEventListener('change',e=>{pageSize=Number(e.target.value);page=1;loadData();});$('nextPageBtn').addEventListener('click',()=>{if(hasNextPage){page+=1;loadData();}});$('refreshBtn').addEventListener('click',loadData);$('jobsRefreshBtn').addEventListener('click',()=>loadJobs());$('startDemoJobBtn').addEventListener('click',startDemoJob);$('sourcesRefreshBtn').addEventListener('click',loadSources);$('sourceForm').addEventListener('submit',saveSource);$('sourceType').addEventListener('change',e=>{$('sourceCsvField').hidden=e.target.value!=='csv';});$('discoveryForm').addEventListener('submit',e=>{e.preventDefault();runDiscovery(true);});$('discoveryRunBtn').addEventListener('click',()=>runDiscovery(false));$('sourcesList').addEventListener('click',e=>{const button=e.target.closest?.('[data-source-action]');if(button)sourceAction(button.dataset.sourceId,button.dataset.sourceAction);});$('addForm').addEventListener('submit',addProspect);$('csvInput').addEventListener('change',e=>{const file=e.target.files[0];if(file){const reader=new FileReader();reader.onload=()=>renderCsv(parseCsv(reader.result));reader.readAsText(file);}});$('menuBtn').addEventListener('click',()=>document.querySelector('.sidebar').classList.toggle('open'));document.querySelectorAll('[data-scroll]').forEach(b=>b.addEventListener('click',()=>document.querySelector(b.dataset.scroll)?.scrollIntoView())); bootstrap(); })(); diff --git a/apps/web/smoke-test.html b/apps/web/smoke-test.html index 87e74b4..647a77c 100644 --- a/apps/web/smoke-test.html +++ b/apps/web/smoke-test.html @@ -42,5 +42,9 @@ frame.onload=async()=>{const d=frame.contentDocument;const js=await fetch('app.j ,['Website scan loading and error states',()=>js.includes('Loading website scan')&&js.includes('Unable to load website scan')&&js.includes('Website scan failed')&&js.includes('role="alert"')] ,['Website scan avoids browser execution claims',()=>!js.includes('screenshot')&&!js.includes('browser execution')&&!js.includes('page screenshot')] ,['Website scan responsive styles',()=>js.includes('website-scan-panel')&&js.includes('website-scan-grid')&&js.includes('@media')] + ,['Public contact extraction smoke marker and controls',()=>!!d.querySelector('[data-smoke="contact-extraction"]')&&!!d.querySelector('#extractContactsBtn')&&!!d.querySelector('#refreshContactExtractionBtn')&&js.includes('/contacts/extract')&&js.includes('jsonRequest')] + ,['Extracted contact evidence fields',()=>['Extracted contacts','Type','Email class','Validation','confidence','Source URL','Suppression'].every(x=>js.includes(x))] + ,['Contact extraction safety and states',()=>js.includes('approved public business pages only')&&js.includes('SMTP')&&js.includes('outreach')&&js.includes('Loading public contacts')&&js.includes('No public contacts found')&&js.includes('Contact extraction failed')] + ,['Contact extraction responsive styles',()=>js.includes('contact-extraction-panel')&&js.includes('extracted-contact-facts')&&js.includes('@media')] ];let passed=0;document.querySelector('#checks').innerHTML=checks.map(([name,test])=>{const ok=test();if(ok)passed++;return `