Regulators block the register, not the rules: what their robots.txt files reveal
Across the bodies we checked, the paths closed to automated crawlers are register lookups and search endpoints. The standards, guidance and policy pages are left open. That pattern tells you which kind of third-party page you can trust and which you must always verify at source.
Primary source: www.teqsa.gov.au · source dated fetched 2026-09-03 · verified on
A robots.txt file is a short public text file at the root of a website that tells automated crawlers which paths they may not visit. It carries no legal force on its own, but it is a deliberate, machine-readable statement of what the site’s operator wants left alone.
We read the robots.txt of every regulator and professional body behind the qualification registers we use. A consistent pattern appeared: the paths they close are register lookups and search endpoints. Their standards, guidance and policy pages are left open.
That pattern is useful to you, because it maps onto a rule for reading anything you find about qualification recognition: explanations of the rules are meant to circulate; lists of who is on a register are not.
What is closed
Checked on 3 September 2026, on the sites’ own robots.txt files:
| Body | Path disallowed to all crawlers | What sits there |
|---|---|---|
| TEQSA | /national-register, /national-register* | The national register of Australian higher education providers |
| CPA Australia | /ManageApplications/AccreditedCourse.mvc/SearchAccreditedCourses* | The accredited course search |
| NZQA | /nzqf/search/*.html$ | Search pages of the New Zealand qualifications framework |
| HCPC | /check-the-register/register-results/ | Results of a register lookup |
| SRA | /search-results/*, /consumers/register/organisation/GetPeople?* | Search results and a register lookup endpoint |
| QAA | /search-results? | Search results |
| AACA | /wp-content/uploads/*.pdf$ | The uploads directory, where its programme list PDF is published |
Two of these deserve a note. CPA Australia’s entry is not a stray line: the same file also blocks its search paths generally, and then names the accredited course search endpoint separately. NZQA’s block is narrow and precise — it targets the .html search pages of the framework specifically, while the rest of the site is open.
What is left open
The same files leave the explanatory material alone. Each of these was reachable on 3 September 2026 and none is covered by a disallow rule on its own site:
- TEQSA’s guides and resources, where its guidance notes on regulation are published.
- HCPC’s standards, the standards of proficiency and conduct that applications are assessed against.
- QAA’s Quality Code, the reference framework for UK higher education quality.
- SRA’s Standards and Regulations, the current rulebook for solicitors in England and Wales.
- The Office for Students blocks only its content management system path, leaving the whole of its published material open.
- Ahpra’s robots.txt blocks its own site search and two script resources, and nothing else.
None of the bodies we checked closes off its rules, its standards, its guidance notes or its policy statements. Several go out of their way to keep them indexable while fencing the register.
The one apparent exception, and why it confirms the pattern
The SRA’s robots.txt also disallows /solicitors/handbook/*. That looks like a rulebook being blocked, which would contradict everything above.
It is not the current rulebook. The SRA Handbook is the regime that preceded the Standards and Regulations, and its contents page still lists the instruments of that earlier regime — among them the QLTS Regulations and the SRA Overseas Rules 2013. The page is still published, and still reachable by a person who follows a link — but it is closed to crawlers, while the rulebook actually in force is left open.
Read that way, the exception says the same thing as the rule, more sharply. What regulators want circulating is the currently operative text. What they do not want circulating is a snapshot that will be mistaken for the current position — whether that snapshot is a copy of the register or a copy of a rulebook that has been replaced.
The same line drawn a different way
Two bodies in our set express the distinction in a newer form, which is worth recognising because it will become more common.
Engineers Australia and the Occupational Therapy Council both publish robots.txt files that disallow a list of named AI crawlers site-wide — among them Anthropic’s ClaudeBot, together with GPTBot, CCBot, Google-Extended, Amazonbot, Applebot-Extended, Bytespider and meta-externalagent. In the same file, both carry the line Content-Signal: search=yes, ai-train=no, use=reference.
Those two instructions are not in tension. Taken together they say: index this site so people can find it, do not ingest it to train models, and referencing it is acceptable. It is a permission to cite paired with a refusal to be absorbed — the same distinction as the register-versus-rules pattern, drawn along a different axis.
A third variant shows how easily this is misread. UK ENIC’s robots.txt reproduces the explanatory comment block that defines what content signals mean, but sets no signal and states no rule. Under that mechanism’s own definition, publishing no signal neither grants nor restricts anything. A reader skimming for the words “content signal” would come away with the opposite impression.
The lesson for reading any of these files is the same: read the whole thing. Directives for named crawlers sit below the general block, and a signal line can qualify everything above it.
Why the two are treated differently
Rules and registers fail in different ways when they go stale, and the bodies behave accordingly.
A rule that is out of date is usually recognisable as out of date. It is dated, it is versioned, it refers to instruments that can be checked, and a reader who lands on a superseded rule can generally tell that something has moved. Rules are also written to be understood and applied by other people; wide circulation serves the regulator’s purpose.
A register entry that is out of date is invisible. A row saying a programme is accredited carries nothing on its face that distinguishes a current entry from one copied three years ago. Accreditation is granted, varied, allowed to lapse and withdrawn. Registers are corrected quietly. A copy captures one moment and then silently misrepresents every moment after it — and the person harmed is the applicant who relied on it.
There is a second reason, which is about authority rather than accuracy. A register is a statement by the body that it, and only it, has made a determination. A convincing third-party copy of that register is a claim to speak for the body. Several of the bodies we read say so directly in their terms of use.
Where the same line is drawn in the terms instead
robots.txt is not the only place a body can make this distinction, and reading only robots.txt will miss the cases where it is made elsewhere.
Ahpra is the clearest example. Its robots.txt blocks its own site search and two script resources, and nothing else — by that signal alone the site is wide open. The restriction on its national registers sits in its terms of use instead, which exclude from the licence it grants “the downloading of any account information, the use of data-gathering or data extraction tools or the downloading and copying of National Register information for any commercial or data storage purpose.”
Same distinction, different instrument. The registers are fenced; the standards, codes and guidance published by the National Boards are not.
This is why the pattern described here is a reading aid rather than a rule you can automate. A body may draw the line in robots.txt, in its terms, in both, or in neither — and where it draws the line in the terms only, a crawler consulting robots.txt would conclude, wrongly, that nothing had been said.
What this means for what you read
Explanations of the rules can be cross-checked, and should be. Because policy, standards and guidance pages are deliberately left open, multiple independent write-ups of the same rule can exist and can be compared. If three sources describe a requirement the same way and each links to the same official page, that agreement is meaningful — and the official page is one click away for you to confirm.
Any third-party list of who is on a register is a copy. It was assembled at some past moment and is not connected to the register afterwards. It receives no corrections and no withdrawals. This holds regardless of how the list is presented, how recently the surrounding page was updated, or how authoritative the site looks. A list without an explicit capture date is worse still, because it withholds the one fact you would need to judge it.
The register is the only thing that answers a register question. Whether a specific provider, programme or practitioner is currently listed is not a question a summary can answer. Nor can it be answered by a page that says it was accurate as of last year. Go to the register and look.
Check the year that applies to you, not just today’s entry. Accreditation takes effect by intake year. What matters for most applications is the position in the year you commenced, which may differ from the position now. Where a register or list offers a dated version or a year filter, use the year you started.
Record what you saw and when. Take the entry, the identifier or course code, and the date. Registers change without notice and rarely publish a change history. A dated record of your own is the only evidence you will have that an entry existed on the day you checked.
A note on how we work
We link to registers. We do not copy them. The pattern described above is one of the reasons: when a body has taken the trouble to close a specific path to automated collection while leaving its rules open, it has expressed a preference clearly enough that there is no need to guess at it.
What we publish is the part that is meant to travel — what the rules say, what the fields on a register mean, where different bodies’ wording conflicts, and how to read an entry once you find it. The entry itself stays where its owner put it.
degree.help summarises publicly available rules and does not replace a formal assessment by a national recognition agency or a professional regulator.
Sources
- CPA Australia — robots.txt (accredited course search endpoint disallowed by name) · fetched 2026-09-03
- NZQA — robots.txt (NZQF search pages disallowed) · fetched 2026-09-03
- HCPC — robots.txt (register results disallowed) · fetched 2026-09-03
- SRA — robots.txt (search results, register lookup endpoint and the superseded Handbook disallowed) · fetched 2026-09-03
- QAA — robots.txt (search results disallowed) · fetched 2026-09-03
- AACA — robots.txt (PDF, XLS and DOC files under uploads disallowed) · fetched 2026-09-03
- Office for Students — robots.txt (only the CMS path disallowed) · fetched 2026-09-03
- Ahpra — robots.txt (site search only) · fetched 2026-09-03
- TEQSA — Guides and resources (open to crawlers; reachable) · fetched 2026-09-03
- HCPC — Standards (open to crawlers; reachable) · fetched 2026-09-03
- QAA — The Quality Code (open to crawlers; reachable) · fetched 2026-09-03
- SRA — Standards and Regulations, the current rulebook (open to crawlers; reachable) · fetched 2026-09-03
- SRA Handbook, the superseded rulebook (disallowed in robots.txt; still published) · fetched 2026-09-03
- Ahpra — Terms (register restriction stated in the terms of use rather than in robots.txt) · fetched 2026-09-03
- Engineers Australia — robots.txt (named AI crawlers disallowed; Content-Signal search=yes, ai-train=no, use=reference) · fetched 2026-09-03
- Occupational Therapy Council — robots.txt (named AI crawlers disallowed; same Content-Signal) · fetched 2026-09-03
- UK ENIC / Ecctis — robots.txt (Content-Signal vocabulary published as comments with no signal set) · fetched 2026-09-03
degree.help summarises published rules. It is not an accreditation body and does not provide immigration advice. Only the named regulator can assess your qualification.