Skip to content

Uploading source material

Source material is the most sensitive thing you hand over. It is your unpublished internal documentation, and it is treated accordingly.

PDF, plain text and Markdown files. Material is uploaded as a file. There is nowhere to give the platform a web address and have it read the page, so if the material lives on your site, export or paste it into a file first.

Up to 40MB per file. The upload screen refuses a larger file before it is sent, so you are told immediately rather than after the transfer, but it is worth knowing before you go looking for material rather than at the moment you upload it. A long text document never approaches 40MB; scans and image-heavy decks routinely do. If a document is over the limit, split it: several files can back one course, and passages from all of them are retrieved together.

Slide decks work well once they are exported to PDF.

Scans are read, not recognised. A PDF is read through its text layer. Nothing performs OCR on the way in, so an image-only scan (pages that are pictures of text) yields no text at all, and the upload is refused with “it is images rather than text” rather than producing a thin course. This is not a quality caveat; it is a wall.

Most scanners and document managers write a text layer as they scan, and a scan that has one is read normally, and that is the case where quality does vary with the scan. If you are not sure which kind you have, open the PDF and try to select a sentence with the cursor. If nothing highlights, run the file through your own OCR tool first and upload the output.

Yours. Your internal training decks, standard operating procedures, manuals, handbooks you authored, your own course notes, your own website content: material your organisation created or commissioned.

By arrangement: material you hold a licence or written permission for, or that is public domain or openly licensed for derivative works. The specific licence must be named, and this route is enabled per account rather than being open by default. Ask us.

Not accepted, regardless of what anyone confirms:

  • Published textbooks and courses you do not own.
  • Examination and certification material: live or recent examination papers, question banks, item banks and marking schemes belonging to any examination board, university or certification body.
  • Anything bearing a confidentiality, “not for distribution” or examination-security marking.
  • Material that is unlawful independently of copyright.
  • Documents whose purpose is personal data: HR files, customer lists, medical records.

Refusing examination content is not refusing the exam-preparation market. A preparation course built from your own training material, or from a published exam-content outline that the certification body itself distributes, is exactly what this is for. The item bank is not.

An automated screen runs twice on every upload: once on the title and file name before a byte is stored, and again on the opening pages once they have been extracted, because a document announces what it is on its cover rather than in its filename. It refuses

  • material described as pirated, leaked or obtained without permission;
  • material carrying a confidentiality or examination-security marking;
  • an examination term sitting beside the name of an examination board or certification body, so “CBSE question paper” but not “question bank” on its own.

An assessment term on its own is not refused. You are asked to confirm it is your organisation’s own material, and that answer is recorded: a company’s own item bank is legitimate and a blocklist would refuse the customers this is built for. Material uploaded on a licence, permission or public-domain basis is refused unless that route has been opened on your account.

It is a screen, not a classifier. It catches the declared and the careless. It cannot recognise a published textbook that somebody has renamed, and it does not look for personal data at all, so the last two rules in the list above are rules you keep, not checks the platform runs. What covers the rest is the rights confirmation below, recorded against the file’s checksum.

Before anything is generated you state, for each file, the basis on which you are entitled to use it:

BasisMeaning
Own workYou or your organisation created it, or it was made for you
LicensedYou hold a licence permitting adaptation and commercial reuse
Public domainCopyright has expired or never applied
Open licencePublished under a licence permitting derivative works
PermissionYou have written permission from the rights holder

The last three, and licensed material, require you to name the specific licence or permission.

Why a basis rather than a tick-box. “I confirm I have the rights” is a box people tick without reading. Naming which basis makes you state a fact about the world. It is a better prompt to actually think, and a far more useful record if it is ever tested.

What is recorded: the file, its checksum, the basis, who accepted, which organisation, the time, and the exact wording that was accepted. The wording is stored verbatim, so that if it is ever revised, old records still show what was actually agreed at the time.

If you replace the file, confirm again. The confirmation is tied to the checksum of the bytes. Different bytes, no valid confirmation, no generation.

It is extracted and split into passages. Each passage keeps its page number and its heading, which is what lets any sentence in a finished course be traced back to the page of yours that it came from, and what makes responding to a complaint a query rather than an investigation.

Two rules, both enforced rather than encouraged:

The syllabus comes from learning outcomes, not from your document. Passages are retrieved to teach each outcome. Your document’s chapter order is never an input. A handbook is ordered for looking things up; a course has to be ordered for someone meeting the material for the first time.

Lessons are written, not copied. Every generated script is compared against the source it was grounded on, on the longest run of identical words and on the proportion of overlapping wording. Anything over threshold is regenerated rather than shipped.

That check finds copied sentences. It cannot tell whether you were entitled to use the material in the first place, and only you know that, which is why the confirmation above exists.

Two different things are described below, and the difference matters more than the numbers.

Deletion on request is immediate, with one exception. An administrator on your side can delete uploaded material in the workspace at any time, and the file and every passage extracted from it go at once. The exception is the one case where the platform refuses: while a rights complaint about that material is open, or once the material has been taken down, the deletion is refused until the complaint is decided. That refusal is deliberate and it is the only thing that holds this path. Without it the party a complaint is about could delete the material, upload the identical bytes, and leave the complaint pointing at nothing.

What survives an ordinary deletion is the record that the material existed: its title, its checksum, the courses built from it and your rights confirmation. That is deliberate too, and it is what makes a later rights complaint a query rather than an investigation. The binding statement of the exception is section 5 of the privacy notice, and clause 13.1 of the data processing agreement says the same thing contractually.

The deletion periods say when something becomes eligible to go, not the second it goes. A period elapses and the material goes on the next nightly run, within a day of its date, not on the stroke of it. The last two rows are the opposite case: nothing deletes them on a date at all, and they are kept on purpose.

WhatKept
Your uploaded fileNo longer than 30 days after the last course is generated from it
An upload no course was ever built fromNo longer than 90 days from upload
Extracted passagesNo longer than 90 days after the course built from them is approved, or after it is published, if that is later
Provenance (source, page, heading, checksum)Life of the course. Deliberately not on a timer
Your rights confirmation7 years, and nothing expires it. See below

A scheduled job, nightly at 05:00 UTC, against the live system. It was not always so: until 1 August 2026 the only thing enforcing these periods was an operator remembering to run the purge by hand, and this page said so. It is now a scheduled job, which is why the rows above are dates rather than intentions.

One thing still delays it, and it is deliberate: the purge refuses to delete anything on a night when no verified database backup was stamped in the previous 24 hours. A failed backup stops the deletions rather than being outrun by them. So a period is a ceiling on a bad night, not a stopwatch: material can sit past its date until a night when the run proceeds.

The privacy notice is the binding statement of all of this: it is section 5 there, and the data processing agreement incorporates that section by reference. This page is the plain-language version of it and must not be read as replacing it.

The seven-year periods have the gap in the other direction: nothing deletes a rights confirmation when its seven years are up. It is kept on purpose, and no code yet expires it.

The evidence deliberately outlives the exposure. A complete copy of your document is the riskiest thing held and is needed only while courses are being produced from it. The record of what you confirmed is what a dispute would turn on years later.

Affected courses are disabled while it is looked into, within 36 hours of a compliant notice, and you are told at the same time as the disablement, with a copy of the notice. That is the commitment the contract makes, so a notice that arrives on a Monday evening can be acted on and answered on the Wednesday morning and still be inside it.

Disablement is not an admission. It is a holding position, and you have a route to contest it. The full procedure is on the terms page.