Vulnerabilities in Internet-exposed web applications are among the most concrete threats to Internet security: once a flaw is disclosed, attackers rapidly exploit exposed instances at scale, provided they can efficiently enumerate them. Certificate Transparency (CT) logs, public and append-only by design, offer such a reconnaissance primitive. In this paper, we characterise this phenomenon from two perspectives: an attacker searching exposed instances of self-hosted web applications, and a defender receiving traffic on exposed services. On the attacker side, we filter a single day of CT logs with simple regular expressions targeting 27 widely deployed applications, yielding over 96000 candidate domains. Despite its simplicity, this methodology proves highly effective: crawling reveals a median match rate of 10.9% (above 20% for eight applications), substantially outperforming baseline enumeration strategies. On the defender side, we deploy 40 honeypots spanning four fidelity tiers across 10 web applications, announcing them in CT logs through TLS certificates for over five months. Probing begins on the day of certificate publication, confirming that attackers actively rely on CT logs. Domain-name keywords have a moderate impact on the traffic received, while realistic honeypots improve visibility and engagement, attracting deeper probing. Strikingly, over 97% of the traffic comes from crawlers operated by AI companies, hinting at an ongoing tectonic shift in Internet background radiation.
Public by Design, Exposed by Default: Web Application Reconnaissance via CT Logs
Drago, Idilio;
2026-01-01
Abstract
Vulnerabilities in Internet-exposed web applications are among the most concrete threats to Internet security: once a flaw is disclosed, attackers rapidly exploit exposed instances at scale, provided they can efficiently enumerate them. Certificate Transparency (CT) logs, public and append-only by design, offer such a reconnaissance primitive. In this paper, we characterise this phenomenon from two perspectives: an attacker searching exposed instances of self-hosted web applications, and a defender receiving traffic on exposed services. On the attacker side, we filter a single day of CT logs with simple regular expressions targeting 27 widely deployed applications, yielding over 96000 candidate domains. Despite its simplicity, this methodology proves highly effective: crawling reveals a median match rate of 10.9% (above 20% for eight applications), substantially outperforming baseline enumeration strategies. On the defender side, we deploy 40 honeypots spanning four fidelity tiers across 10 web applications, announcing them in CT logs through TLS certificates for over five months. Probing begins on the day of certificate publication, confirming that attackers actively rely on CT logs. Domain-name keywords have a moderate impact on the traffic received, while realistic honeypots improve visibility and engagement, attracting deeper probing. Strikingly, over 97% of the traffic comes from crawlers operated by AI companies, hinting at an ongoing tectonic shift in Internet background radiation.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_WTMC_CT_Logs_AAM.pdf
Accesso aperto
Descrizione: Accepted manuscript (camera-ready)
Tipo di file:
POSTPRINT (VERSIONE FINALE DELL’AUTORE)
Dimensione
750.93 kB
Formato
Adobe PDF
|
750.93 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



