Immigration platform + IRCC autofiller
Bilingual intake for an immigration consultancy: a conditional assessment, structured client data, and a Chrome extension that fills the government portal.
Context
From November 2024 to August 2025 I worked with a regulated Canadian immigration consultancy: one RCIC-licensed consultant and their staff, serving Arabic-speaking applicants. A development team built the website and the admin dashboard. I wrote the specification that directed their work, designed the data model and the Free Assessment, and built the form-filling pipeline myself.
Problem
The last step of an application is IRCC’s online portal. For EMPP, its digital forms are IMM 0008, IMM 5406, IMM 5669 and IMM 5562, with pages repeated for the spouse and the children. Client details used to arrive in email chains, and staff transcribed them by hand into PDF forms and the portal.
- Sensitive data. Passport and national ID numbers, ten years of addresses, criminality and health questions.
- Two languages. Every client-facing string in English and Arabic.
- IRCC owns the questions. The data has to map to the government’s fields one by one, and those fields change.
What I built
The Free Assessment. I designed it as a bilingual decision tree; on the site it runs as a multi-step form. It collects contact details, then branches. Applying from inside Canada leads to eight services, from Super Visa to PR renewal. Otherwise, questions on work, family, IELTS scores, funds and refugee status lead toward Express Entry, OINP, LMIA, EMPP, private sponsorship (SAH), or a visitor, business or study visa. The tree encodes concrete thresholds: IELTS 8/7/7/7 with a two-year diploma, a bank balance above USD 20,000, IRCC’s proof-of-funds table by family size. Every result ends in a Calendly booking link, except one: “no programs available at this time”. Answers are stored through the API.
A specification and a data model. I wrote the platform spec in five parts, plus a database design. The model mirrors the forms. Each IRCC question becomes one short key: “Passport number” is passport_num, “Date of birth (YYYY-MM-DD)” is birth_date. One person table holds everyone on an application, tagged by relationship. Addresses, education, travel and criminality answers hang off each person. The intake covers 18 people: the principal applicant and the spouse, each with up to three children, two parents and up to three siblings.
In the spec, staff assign a package, such as EMPP, and the client portal shows only those forms. For the review side I prototyped a data viewer: one card per person, a modal with everything on file, and click-to-copy on every value. The spec keeps it read-only; only an admin can toggle editing.
The spec routes static site text through a Google Sheet: 109 keys in 11 sections, English and Arabic side by side. A developer runs a script that turns it into en.json and ar.json, so copy changes need no code.
An LLM pipeline over the government’s forms. I saved a complete local copy of the portal’s forms: 65 pages: the portal’s two main pages and the four IMM forms. A Python script stripped each page down to its form markup, dropping scripts, styles and session-timeout dialogs. An LLM read the 60 stripped form pages and returned every field as structured data: label, name and id, type, options, required flag, help text, and repeating groups like children or trips. A second script kept the principal applicant’s 217 text fields as JSONL, one line per field: section, key, label, value. Family members’ pages reuse most of the principal applicant’s field names (190 of the spouse’s 231), so the same format works for each person. The model did the reading; every key is the portal’s own name attribute.
A Chrome extension. My first plan proposed a Tampermonkey userscript. v1, packaged in June 2025, is a Manifest V3 Chrome extension for staff. Load a person’s JSONL file, open a portal page, press Fill Form. It finds each field by name, writes the value, and fires input and change events so the portal’s form code registers it. Then it shows “Please review the form” and stops. There is no submit step.
How I knew it worked
Offline first. The local copy let me build and test without the live portal. Each page works on its own; only the links between pages are dead. Three sample applicant files exercised the fill.
Real keys. All 170 distinct keys in the JSONL appear as name attributes in the 60 stripped pages. The model transcribed them and invented none.
Status at handover. In August 2025 I handed over the spec, the scripts, the sample data, an issue list and a 36-minute walkthrough video. The spec records where each part stood. The Free Assessment worked and needed no changes. The services and news CMS was complete. The extension was a functional proof of concept. The client portal was still in development.
v1 was delivered with four known bugs, all in the issue list. Passport and national ID dates got mixed up. So did everyone’s education dates. The spouse’s trip destination wouldn’t fill. Dates in general were unreliable.
The first two share one cause, and I built it in: fields keyed by the name attribute alone. My plan called name stable and “unique per field”. It is stable, not unique. The passport and national ID pages both use issueDate and expiryDate. Five IMM 5669 pages, from education to addresses, all use from0 and to0. In the JSONL, 27 keys repeat across 74 of the 217 rows. The fill loop writes every row whose name matches, so the last one wins.
v1 also can’t pull from the platform. Each person’s data has to be prepared as a JSONL file and loaded by hand. That gap is why the v2 design exists: staff sign in with dashboard credentials, pick an application from the API, and fill from live data. It is designed, not built.
What I’d change
Make keys unique, and check it. Form, page and name, not name alone. The converter should refuse a duplicate key instead of letting it surface as a wrong date in the portal.
Check the extraction before trusting it. The model found a label for every field except dates: all 64 date fields came back “Label not found”. A reviewer can’t verify a value they can’t identify. I’d treat an unlabelled field as an extraction failure, not as data.
Carry every field type. The converter kept text inputs only: 217 of the 284 fields on the principal applicant’s pages. The fill code already handled selects, radios and checkboxes; the data files never carried them.
Links
None public. The client is anonymized and the platform is private.