JurisDATA·2019–2022

Tutorial: build a safe public-record intake workflow

How a public legal guide became a 50,000-person intake pipeline: explicit retrieval outcomes, duplicate-safe records, approved routing, and channel-aware delivery.

Open public reference →

This is a technical tutorial from JurisDATA, a legal-information project I built with Alice, a lawyer. It describes an operational system, not legal advice. The system helped people move from a public guide about photo-enforced traffic tickets to a careful, case-specific next step.

Start with the decision boundary

The article ranked for relevant searches and brought a large volume of people asking the same question: what does my record show, and where do I start? A browser check and manual follow-up for every request did not scale. An automatic legal conclusion would have been unsafe. The useful boundary was narrower: retrieve available public information, retain the result with the right person, and route only approved guidance.

The workflow, as a contract

The pipeline had six stages: intake; public-record retrieval; structured outcome; duplicate-safe person and record storage; approved classification; and WhatsApp or email delivery. Each stage produced data for the next stage, rather than an unstructured message that later systems had to interpret again.

The implementation connected Make, Bright Data for SIMIT retrieval, Airtable, Fluent CRM, Hookdeck, Twilio, and email. Tool names were secondary. The durable design was the contract at each handoff: one intake maps to one known person, one retrieval outcome, and one permitted next-step route.

Screenshot: end-to-end intake map

Add a sanitised workflow screenshot here. Show only: article or form → retrieval → structured result → CRM record → approved route → WhatsApp or email. Redact all identifiers, request fields, credentials, and live records.

Model retrieval results explicitly

A returned record and a no-record result are both valid results. They must not collapse into the same empty value. The workflow kept the result state explicit so that a missing result did not look like a confirmed answer and an incomplete response could be reviewed instead of silently routed.

This is the pattern to reuse in any public-data intake: represent FOUND, NO_RECORD, and NEEDS_REVIEW as different states; keep the original request connected to the result; and permit downstream communication only for states that have an approved route.

Make retries duplicate-safe

The system looked up an existing person and an existing record before creating another one. This matters because webhook delivery, browser retrieval, and CRM writes can all be retried. Without that lookup, a normal retry can become duplicate contacts, duplicate records, and conflicting follow-up.

The practical rule is simple: decide which stable facts identify the person and the retrieved record, perform the lookup before the create, and treat an existing record as a recovery path. Do not use a delivery timestamp as the identity of a request.

Separate classification from legal judgment

Automation can classify structured operational facts. It should not manufacture legal advice from incomplete information. The workflow used defined routes and approved message content. The legal criteria and content remained owned by the people responsible for them; the automation moved the appropriate facts to the appropriate route.

Screenshot: classification rule

Add a redacted Airtable formula or a redrawn decision table here. Show the allowed outcome categories, not personal data, case values, or unpublished legal criteria.

Treat delivery as its own reliability problem

A second workflow received an approved classification, checked that a usable phone number existed, normalised it for Colombian delivery, and sent the matching message through Twilio. A missing or unusable number was not a reason to send malformed data. It was a delivery state that needed a different route or review.

This separation keeps retrieval failures, classification decisions, and channel failures visible. It also makes message delivery replaceable: the record and approved route survive if a channel changes later.

Reusable checklist

For a high-volume public-information workflow: define outcome states before selecting tools; keep requests, people, and records linked; perform duplicate checks before creates; make approved routes explicit; validate channel data at the delivery boundary; and preserve human review where the system lacks evidence. Over time, this design supported more than 50,000 people without pretending automation could replace legal judgment.

Publication boundary

The screenshots on this page show the system as it operated at the time. They must not expose personal data, credentials, unpublished decision criteria, or live public-record responses. This tutorial is not legal advice.