---
title: "Should You Build Email Validation In-House?"
description: "An honest build-vs-buy answer for email validation: the MX-checking core is an afternoon, and the five behaviours that make it correct — typo-before-DNS, the SERVFAIL/NXDOMAIN split, typosquat hosts, null MX, cache — are the recurring cost."
slug: "build-email-validation-in-house"
date: 2026-10-04
updated: 2026-10-04
last_tested: 2026-10-04
summary: "You can write a working MX validator in twenty lines today. This page enumerates what that version silently gets wrong against live DNS, and states plainly when building it yourself beats paying for an API — and when it does not."
cluster: "Choosing a validator"
intent: decision
sources:
  - title: "RFC 5321 — Simple Mail Transfer Protocol (implicit A/AAAA fallback, section 5.1)"
    url: "https://www.rfc-editor.org/rfc/rfc5321.html"
  - title: "RFC 7505 — A Null MX Resource Record for Domains That Accept No Mail"
    url: "https://www.rfc-editor.org/rfc/rfc7505.html"
  - title: "RFC 1035 — Domain Names: Implementation and Specification (RCODEs)"
    url: "https://www.rfc-editor.org/rfc/rfc1035.html"
  - title: "dnspython documentation"
    url: "https://www.dnspython.org/"
  - title: "nobounce.dev verdict taxonomy"
    url: "https://nobounce.dev/v1/config"
  - title: "nobounce.dev agent reference"
    url: "https://nobounce.dev/llms.txt"
---

Short answer: if you need to validate addresses in a batch job against domains you already understand, and you control the DNS resolver, build it — the core is twenty lines of `dnspython`. If this guards a signup form, where the value is catching `user@gmai.com` in the 200 ms before Submit, the recurring cost is not the code but the domain knowledge: typo correction that runs *before* DNS, the SERVFAIL/NXDOMAIN split, typosquat MX hosts, and a cache you never get to warm. That is what you are buying at $1/mo, and it is a legitimate choice to build it anyway. The rest of this page is the evidence for both halves, checked against live DNS behaviour rather than vendor claims.

## The afternoon version, honestly

This works, and it is not a straw man — it is roughly what most in-house validators are:

```python
import dns.resolver

def validate(domain: str) -> str:
    try:
        answers = dns.resolver.resolve(domain, "MX")
    except dns.resolver.NXDOMAIN:
        return "domain_not_found"
    except (dns.resolver.NoAnswer, dns.resolver.NoNameservers):
        return "no_mx"
    if any(r.exchange == "." for r in answers):
        return "no_mx"
    return "ok"
```

Parse the syntax with a regex, call this, reject or accept. For a one-off import against corporate domains, this is often *enough*, and you should not pay anyone $1/mo to avoid twenty lines you understand.

## What the afternoon version gets wrong

Each item below was verified against live DNS, not against documentation. They are the reason the pipeline order and the reason codes exist.

**1. SERVFAIL is not NXDOMAIN.** The code above collapses `dns.resolver.NoNameservers` (SERVFAIL, RCODE 2 per RFC 1035) and "no MX" into one branch. They are opposite answers: NXDOMAIN means the domain does not exist; SERVFAIL means *the resolver could not complete the lookup* — a broken delegation, a timeout, an upstream outage. Rejecting on SERVFAIL means your signup dies whenever DNS hiccups, which is a self-inflicted outage. The correct behaviour is to fail open and report "unknown, DNS did not complete" so you can audit it later. The split is covered in depth in [SERVFAIL vs NXDOMAIN in Email Validation](https://nobounce.dev/blog/servfail-vs-nxdomain-email-validation/).

**2. Null MX and the A/AAAA fallback.** RFC 7505 defined `0 .` so a domain can *publicly* refuse all mail; treating that as "no records, try the A record" misses the point. Meanwhile a domain with no MX but a live A/AAAA record still accepts mail per RFC 5321 section 5.1 — so a naive MX-only check rejects deliverable domains. You need both behaviours, plus the RFC 2606 reserved names (`example.com`, `*.test`) separated out, or your own QA fixtures start failing validation in ways that look like user typos. That trap has its own write-up in [Null MX and Reserved Domains](https://nobounce.dev/blog/null-mx-rfc-7505-reserved-domains/).

**3. Typos resolve.** This is the finding that inverts naive architectures: `gmai.com` and `hotmial.com` have live MX records. A DNS-first pipeline returns "deliverable" for an address the user never meant to type, which is worse than rejecting it — you accept a bounce you could have corrected. Fixing this means Damerau-Levenshtein correction (transpositions like `gmial`→`gmail` must score as one edit) against a curated list of high-volume domains, run *before* any DNS lookup, and someone has to maintain that list. The full evidence is in [MX Validation Alone Misses Typo Domains](https://nobounce.dev/blog/mx-validation-misses-typo-domains/).

**4. Typosquat MX hosts are a dataset, not a constant.** `gmai.com` and `hotmial.com` are both fronted by the same operator host (`mail.h-email.net`). Clustering domains by MX host finds these in bulk — but the list grows as operators rotate infrastructure, so it is a table you update, not a line of code. This is described in [Typosquat MX Host Clustering](https://nobounce.dev/blog/typosquat-mx-host-clustering/).

**5. The cache is the product economics.** Live DNS checks are the only marginal cost in validation. A shared cache across every customer means the common domains are warm hits and the check is effectively instant; a self-hoster starts cold forever, pays the full latency on every lookup, and pays it again after every TTL expiry. A failed lookup must also never be cached, or one outage pins a wrong verdict for the TTL.

None of this is unbuildable. It is all buildable — and then it is *yours*, which means a quarterly chore of domain lists, SERVFAIL post-mortems and cache tuning. The question is whether that chore is your product or a distraction from it.

## What building it yourself does not get you either

Do not build in-house expecting mailbox-level certainty. The twenty-line version, the $1/mo API and the $500/mo enterprise tier all face the same wall: confirming that a specific mailbox exists requires SMTP RCPT TO probing, and the established vendors' claims in that direction are weakened by catch-all domains and by IP reputation games. A `deliverable` verdict — from your code or from any API — means the syntax is right and the domain accepts mail, and nothing more. No honest tool guarantees delivery of the next message.

## The decision

Build it in-house when:

- it validates **your own controlled domains** in tests or internal pipelines — an external API is the wrong dependency there;
- the volume is a one-off or monthly batch where 200 ms latency and typo recovery do not matter;
- you want the DNS learning — it is a genuinely good exercise.

Buy when:

- it sits **inline in a signup flow**, where the recoverable typo is revenue and the suggestion field is the point;
- you need **fail-open semantics you can audit** (an `unknown` verdict with `checked.mx: false`, not a guess);
- **retention matters**: the in-house version logs whatever you log; nobounce persists no address in any form, only SHA-256 hashes, because the shared cache keys on domains.

If that lands you on "buy", do not take this page's word for it — point an agent at the evaluation document and let it exercise the engine against your own cases, including the ones where nobounce is the wrong tool:

```
Evaluate https://nobounce.dev for my project. Fetch and follow:
https://nobounce.dev/evaluate.md
```

If it lands on "build", keep the five items above as your acceptance tests. They are the difference between a validator and a bounce generator.
