A permissive robots.txt is not permission: the two layers to check before you use a source
Four bodies whose robots.txt files are wide open prohibit copying or automated extraction in their terms of use — one of them naming data-gathering and data extraction tools directly. Being able to retrieve a page is not the same as being allowed to use it.
Primary source: www.ahpra.gov.au · source dated fetched 2026-09-03 · verified on
A site’s robots.txt tells automated crawlers which paths they may not fetch. Its terms of use say what you are allowed to do with what you fetched. These are different documents, written by different people, answering different questions — and they routinely disagree.
When they disagree, the terms are the operative statement. robots.txt is a convention with no legal force of its own; terms of use are asserted as a contract. A wide-open robots.txt is not a grant of anything. It is the absence of a technical instruction.
This matters for anyone doing comparison research on qualification recognition, because the bodies holding that data are unusually likely to be in exactly this position.
Four sources where the two layers say different things
All four were checked on 3 September 2026. In each case the robots.txt permits essentially everything, and the terms of use prohibit the use you would want.
Ahpra, which holds Australia’s national registers of health practitioners, has a robots.txt that blocks its own site search and two script resources, and nothing else. Its terms of use state that the limited licence it grants excludes “the downloading of any account information, the use of data-gathering or data extraction tools or the downloading and copying of National Register information for any commercial or data storage purpose.” It adds that no rights are granted to download or modify the registers, other than page caching. This is the case where the prohibition names the tooling directly.
The British Psychological Society has a robots.txt that blocks two preview paths. Its terms of use state: “You must not conduct any systematic or automated data collection activities (including without limitation scraping, data mining, data extraction and data harvesting) on or in relation to our website without our express written consent.”
World Education Services has a robots.txt that blocks three specific PDF files. Its terms and conditions state that “you are prohibited from copying, reproducing, modifying, distributing, displaying, performing, or transmitting any of the contents of the Services for any purposes (other than those expressly authorized in writing by WES).” The phrase to notice is “for any purposes” — there is no exception for research, education or non-commercial use.
The Australian Association of Social Workers has a robots.txt consisting of an empty disallow directive, which means every path is permitted to every crawler. Its terms of use state: “No part of the Information may be copied, reproduced, modified, republished, uploaded, posted, transmitted or distributed in any form or manner without AASW’s prior written consent.”
In all four cases, a script that consulted only robots.txt would conclude the site was open. In all four cases that conclusion would be wrong.
The mirror image also happens
The reverse case is just as common and is often misread in the other direction.
TEQSA’s robots.txt disallows /national-register and /national-register* while leaving the rest of the site open. That is a statement about automated retrieval of one path. It is not, by itself, a copyright restriction — TEQSA separately publishes an open Creative Commons licence covering its website material.
A second variant is worth recognising: the block that targets a file type rather than a query path. AACA’s robots.txt disallows PDF, XLS and DOC files under its uploads directory — which is where a body of that kind publishes its lists. Nothing there speaks to copyright at all; it is an instruction about retrieval, aimed at the artefact rather than at a search endpoint.
So a disallow line does not automatically mean “you have no rights here”, and an allow does not mean “you have rights here”. The two layers genuinely are independent, and each can be silent while the other speaks. You have to read both.
A third layer: when you cannot read the rule at all
Two bodies whose terms we wanted to re-check on 3 September 2026 returned a firewall block page — to our crawler and to an ordinary desktop browser user agent alike — including for robots.txt itself. In those cases there is no readable instruction of any kind.
That situation calls for the conservative reading. If you cannot retrieve a site’s stated position, you do not have permission by default; you have an unknown. The honest move is to record it as unverified and either use a different source or contact the body, rather than to treat silence as consent.
A fourth signal worth knowing about
Some sites now use robots.txt to distinguish between kinds of automated use, rather than to allow or block wholesale.
Engineers Australia and the Occupational Therapy Council both publish robots.txt files that disallow a list of named AI crawlers — including Anthropic’s ClaudeBot, along with GPTBot, CCBot, Google-Extended, Amazonbot, Applebot-Extended, Bytespider and meta-externalagent — while also carrying the line Content-Signal: search=yes, ai-train=no, use=reference.
Read together, those say something specific and coherent: index this site for search, do not use it to train models, and referencing it is fine. That is not a site closing itself off. It is a site drawing a line between being cited and being absorbed.
UK ENIC’s robots.txt shows the same vocabulary used differently: it publishes the explanatory comment block describing what content signals mean, but sets no signal and states no rule. Under that mechanism’s own definition, publishing no signal neither grants nor restricts anything.
The practical point is that robots.txt is no longer a simple yes or no. It can now carry a statement of intent, and reading it properly means reading the whole file rather than pattern-matching on Disallow: /.
The check, in order
Before relying on any source for a research or comparison project:
1. Fetch the robots.txt first, before anything else. It is at https://<domain>/robots.txt, it is plain text, and it costs one request. Read the whole file, not just the User-agent: * block — directives for named crawlers frequently sit below it, and a Content-Signal line may qualify everything above it.
2. Then find the terms of use, copyright page or disclaimer. Look in the footer first, then try /terms, /terms-of-use, /copyright and /disclaimer. Note that a page labelled “Terms” is sometimes a privacy policy or a procurement document; check that what you are reading actually addresses re-use.
3. Look for the specific words that describe what you intend to do. “Republish”, “reproduce”, “distribute”, “systematic or automated data collection”, “scraping”, “data extraction”, “data mining”, “commercial purpose”. A general copyright assertion in a footer is weaker than a clause naming your activity, and a clause naming your activity settles the matter.
4. Check whether any permission granted is limited to non-commercial use. This is the most common trap. Several bodies grant re-use only for personal, educational or non-commercial purposes. If your project carries advertising, affiliate links or any commercial character, that permission does not reach you, and reading it as though it does is wishful thinking rather than research.
5. Treat silence as no. If neither the robots.txt nor the terms addresses re-use, the default position under copyright is that rights are reserved. A footer reading ”© 2026” or “all rights reserved” asserts rights; it does not grant them. Silence means the question is open, and the safe resolution of an open question is to link rather than copy.
6. Record the date and the wording. Terms pages are revised without announcement. Note what the page said and when you read it, so a later change does not quietly undermine what you relied on.
7. Where a body offers an application route, that is the answer. Several publish an address or a form for permission requests, and some operate a paid licence. If the data genuinely matters to what you are building, asking is available and is the only route that produces a durable answer.
Why this is worth your time as a reader, not just as a builder
You may never scrape anything. The two-layer check still tells you something about the pages you read.
A third-party site presenting a full copy of an official register has, in most cases, taken data that the holding body did not license for that use. That does not automatically make the copy inaccurate. It does tell you the copy exists outside any relationship with the source: it receives no corrections, no withdrawals and no notice of change, because there is no channel through which those could arrive.
So the two layers give you a reading habit. When a site shows you a rule and links to the official page it came from, you can check it and the check is cheap. When a site shows you a register, ask where the data came from, when it was captured, and whether the source permits it to be there. If those questions have no answer on the page, use the register itself.
degree.help summarises publicly available rules and explains what they mean. It does not assess qualifications and does not replace a formal evaluation by a national recognition agency or a professional regulator. Nothing here is legal advice on copyright, licensing or website terms; the licence text and the body that published it are the authorities on their own material.
Sources
- Ahpra — robots.txt (site search and two script resources only) · fetched 2026-09-03
- BPS — Terms of use (systematic or automated data collection prohibited) · fetched 2026-09-03
- BPS — robots.txt (preview paths only) · fetched 2026-09-03
- WES — Terms and conditions (copying prohibited for any purposes) · fetched 2026-09-03
- WES — robots.txt (three PDF files only) · fetched 2026-09-03
- AASW — Terms and conditions (no part of the information may be copied or republished) · fetched 2026-09-03
- AASW — robots.txt (empty disallow; everything permitted) · fetched 2026-09-03
- Engineers Australia — robots.txt (named AI crawlers disallowed; Content-Signal search=yes, ai-train=no, use=reference) · fetched 2026-09-03
- Occupational Therapy Council — robots.txt (named AI crawlers disallowed; same Content-Signal) · fetched 2026-09-03
- UK ENIC / Ecctis — robots.txt (Content-Signal vocabulary published as comments with no signal set) · fetched 2026-09-03
- TEQSA — robots.txt (National Register disallowed while the rest of the site is open) · fetched 2026-09-03
- TEQSA — Copyright (Creative Commons Attribution 3.0 Australia over site material) · page states Last updated 13 Oct 2022; fetched 2026-09-03
- AACA — robots.txt (PDF, XLS and DOC files under uploads disallowed) · fetched 2026-09-03
degree.help summarises published rules. It is not an accreditation body and does not provide immigration advice. Only the named regulator can assess your qualification.