Ask a vendor about employee records and they will show you search.
Natural-language queries across personnel files, instant retrieval of a contract signed four years ago, a chat interface that answers “when did she change job title”. It demonstrates beautifully, and finding things was never what made employee records expensive.
What makes them expensive is the opposite operation. Knowing what you hold. Knowing why you are still holding it. Being able to remove it when someone asks, completely, from every place it ended up. Search is a convenience. Deletion is a liability, and it is the capability that cannot be bolted on afterwards.
Three Questions Before You Look at Any Product
These take an afternoon and they eliminate most of the shortlist.
Where do employee records currently live? The honest answer in a small business is never one system. Contracts in a document store, payroll data in a payroll provider, performance notes in a manager’s mailbox, right-to-work scans in a shared folder, disciplinary records somewhere deliberately obscure. A tool that indexes one of those and calls itself a records system has solved a fraction of the problem while making the remainder harder to see.
What are you required to keep, and for how long? Retention periods differ by record type and jurisdiction, and most small employers keep everything forever because that felt safer than deciding. It is not safer. Data you are not required to hold is pure liability in a breach and pure cost in a subject access request.
Who can see what? Personnel files are the most sensitive records most small businesses hold, and access is usually controlled by folder permissions somebody set in a hurry. A records tool inherits whatever access model you give it, and an AI layer that can answer questions across everything will happily answer them for whoever asks.
Our piece on AI in data management covers the general shape of getting scattered data under control.
Where AI Genuinely Helps
Three capabilities are worth paying for, and none of them is the chat interface.
Classification of the Backlog
Every organisation has a folder nobody has opened since a migration. Thousands of documents, no consistent naming, some of them personnel records and some of them not.
Classification is a genuinely good fit for the technology. A model reading each document and tagging it as a contract, a payslip, a medical note, a reference or a disciplinary record does in a week what nobody was ever going to do by hand. It is high volume, mechanical, and forgiving of the occasional error because a human reviews the categories rather than each item.
The limit is confidence. A classifier that labels everything with equal certainty will file a handful of sensitive documents in a general bucket, and those will be the ones that matter. Insist on confidence scores and review the low-confidence tail by hand, since that tail is small and contains most of the mistakes.
Retention Dates Attached at Ingestion
The reason retention is never enforced is that it depends on knowing a document’s type and its trigger date, and that information exists only in somebody’s head at the moment of filing.
A tool that assigns a retention date when a record arrives, derived from its type and the relevant date, converts a policy nobody follows into a field on a record. That is the whole trick, and it is quietly transformative.
What to check is what happens when the date arrives. A tool that deletes silently is a compliance problem of a different kind, and a tool that generates an alert nobody owns will accumulate expired records with alerts attached, which is worse than no alert because it documents that you knew. The workable middle is a review queue with a named owner and a bulk approve.
Assembling a Subject Access Request
When a current or former employee asks for their data, the clock starts and the work is enormous: everything held about one person, across every system, redacted where it names others.
An AI tool that assembles the candidate set in hours instead of weeks is worth the licence on its own. The redaction it proposes still needs review, because deciding what constitutes a third party’s information is a judgement call, but assembly is the part that consumes the deadline.
Ask specifically which systems it can reach. A tool that searches the document store and not the mailboxes has done the easy half, and mailboxes are where the difficult material usually is. Our article on Office 365 management covers the mailbox side of that estate.
Where It Should Not Be Trusted
Two categories, and both are about consequence rather than capability.
Anything that feeds an employment decision. A model summarising a performance history for a manager considering redundancy is producing evidence in a process that may be examined by a tribunal, and a summary that drops context is indistinguishable from one that does not until somebody checks. Managers should read the record.
Automatic deletion without review. The scenario is specific: a retention rule fires and removes a document that has become relevant to an active dispute. Legal hold has to override retention, which means the tool needs to know a hold exists, which means somebody has to be able to set one in seconds. Check that this exists before you enable any automated deletion at all.
The Migration Is the Project
The tooling decision matters less than most vendors suggest. The migration decides whether you get value, and it is consistently underestimated.
Three things go wrong. Documents move without their metadata, so the retention dates you were buying the tool for are absent on everything historical. Duplicates arrive as separate records, so one employee has three overlapping files and no version is authoritative. And the old locations stay populated, because deleting the source felt risky during a migration, which means you now hold everything twice and the tool’s picture of what you hold is wrong from day one.
That third one is the most common and the most damaging, since it silently defeats both retention and deletion. Plan the source decommission as part of the migration rather than a task afterwards, and give it a date.
Records that live on personal or mobile devices deserve a specific mention, since they are routinely outside whatever system you buy. Our mobile device management work exists partly for this reason.
What Adoption Actually Requires
A records system fails quietly when people keep working around it.
The pattern is familiar. A manager saves a performance note to their own drive because it is faster, an offer letter is emailed rather than filed, a signed variation sits in a mailbox. Six months on the system holds the formal records and the mailboxes hold everything that matters, and nobody notices until a subject access request arrives.
What prevents it is making the filing path the shortest path rather than the compliant one. Filing from the mailbox in one click, from the phone, from wherever the document arrives. If filing takes more effort than not filing, people will not file, and no amount of policy changes that arithmetic. This is the same behavioural point our piece on using technology to boost employee efficiency makes about tooling generally.
Deciding which of your existing systems can even support this is ordinary AI readiness work, and the ingestion and classification pipeline itself is intelligent process automation applied to an HR function. The wider organisational shift is covered in our article on AI transforming employee management.
Leavers, Which Is Where Most Records Estates Go Wrong
The moment an employee leaves is when a records estate either holds together or quietly fractures, and almost no small business has a defined process for it.
What typically happens is that the mailbox is converted to a shared one so a colleague can pick up client traffic, the personal drive is left in place because somebody might need something, the HRIS record is marked inactive, and the physical file goes into a cabinet. Four different fates for one person’s records, none of them recorded anywhere.
The consequence surfaces years later. Somebody asks what you hold about a former employee, and the honest answer requires searching a shared mailbox that has since accumulated three other people’s correspondence, a drive nobody has opened, and a cabinet. The subject access clock does not care that the estate was tidy when they left.
What works is a leaver checklist that treats records as a named step rather than an afterthought, with one decision per location: transferred, retained until a stated date, or deleted now. Any tool worth buying can hold that decision against the person, so the answer exists in a system rather than in the memory of whoever handled the exit.
The Shared Mailbox Problem Specifically
This one deserves separating out, because it is nearly universal and it is a genuine conflict.
Converting a leaver’s mailbox to shared is operationally sensible: client threads continue, nothing is lost, the business keeps running. It is also the single most common way personal data outlives its retention period, because a shared mailbox is nobody’s responsibility and never gets reviewed.
There is no clean answer, only a less bad one. Set a date at conversion, and at that date either the mailbox is archived under a retention rule or it is deleted. The date matters more than which option you pick, because the failure mode is not choosing wrongly, it is never revisiting the decision at all.
How We Score Employee Records Tools
Four questions, in this order.
Can it tell you everything it holds about one named person, across every connected system, in one operation. Does it attach a retention date at ingestion, derived from record type rather than typed by hand. Can a legal hold be applied in seconds and does it override retention. And is the access model per record type rather than per folder, so a manager can see contracts without seeing medical notes.
Search quality is not on the list. Every serious product searches well, and none of them fails on it.
Frequently Asked Questions
Do we need a dedicated records system or is our HRIS enough?
If your HRIS holds documents with retention dates and can produce everything about one person on request, it is probably enough. Most cannot do the second thing, which is what pushes small employers toward a dedicated document layer alongside it.
What should we fix before migrating anything?
Decide retention periods by record type, and agree what happens to the source locations after the move. Migrating without either means you arrive with no retention dates and two copies of everything, which is worse than where you started.
Can AI decide what to delete?
It can propose, based on the retention rule and the record type, and a person should approve the batch. Automatic deletion without a legal-hold override is the one configuration we would not run, because the failure destroys evidence in an active matter.
How do we handle records in mailboxes?
Treat them as in scope, because a subject access request does. Either bring them into the records system at the point they arrive, or make sure whatever you buy can search them. A tool that cannot reach mailboxes has solved the easy half.
Who should own the system day to day?
One named person, usually in HR rather than IT, with a deputy. Records questions are judgement calls about sensitivity and retention rather than technical ones, and a system owned by nobody drifts back to folders within a year.
Who Is Behind This Advice
Our team has run these migrations at businesses where the HR function is one person who also does office management. That shapes the advice: the sophisticated retention matrix is the one that gets abandoned, and what survives is four record types with four periods, applied consistently.
We have also seen the failure this article is written against. A client bought a capable platform, indexed the document store, left three years of personnel material in the old shared drive because deleting it felt risky, and then received a subject access request. The tool answered confidently and incompletely, which was the worst of both outcomes.
Matt Rosenthal leads Mindcore and pushes for records systems that reduce what you hold rather than simply organising it, which is a harder sell and a much better position to be in.
Buy for Deletion, Not for Search
Settle this before you shortlist. Write down where employee records currently live, including the mailboxes and the personal drives, because that list is the actual scope. Agree retention periods by record type, keeping the matrix small enough to be applied consistently. Require retention dates to be attached at ingestion and derived from type, not typed in. Confirm a legal hold can be set in seconds and overrides retention, and do not enable automated deletion until it can. Insist on per-record-type access rather than folder permissions. Check the tool can search mailboxes, not only the document store. Plan the source decommission with a date, as part of the migration rather than after it. And make filing the shortest path from wherever a document arrives, since a system people work around holds only the records nobody needed. If you want help mapping where your employee records actually sit before you buy anything, book a free strategy call and we will start with that list.


