Audit evidence management is the process of collecting, organizing, retaining, and presenting the artifacts that prove a control operated as designed. Under AU-C 500 and PCAOB AS 1105, that evidence must be both sufficient (enough of it, driven by risk and control frequency) and appropriate (relevant and reliable by source). Inquiry alone can never prove a control's operating effectiveness — you need system-generated records.
Key takeaways
- Sufficient and appropriate are two different tests. Sufficiency is quantity (how many samples); appropriateness is quality (relevance + reliability). Both must be met — AU-C 500 and AS 1105.
- Reliability follows a hierarchy. System-generated reports, immutable logs, and reperformed results outrank screenshots and inquiry; auditor-obtained evidence outranks client-asserted evidence.
- Sample size scales with control frequency, not report type. The AICPA mandates no fixed number; auditors size samples under AU-C 530 and the AICPA Audit Sampling Guide, targeting roughly 90% confidence for larger populations.
- Retention differs by framework. HIPAA fixes six years (45 CFR 164.316(b)(2)); GDPR sets no number (storage-limitation principle); the CPA firm keeps its own workpapers at least five years (AU-C 230).
What is audit evidence management?
Audit evidence management is the discipline of collecting, organizing, retaining, and presenting the artifacts that prove a control operated as designed. In an assurance engagement, the auditor does not take your word for it — they form an opinion by gathering evidence and testing it. Your job as the audited organization is to make that evidence complete, attributable, timely, and easy to inspect.
Two words govern everything an auditor does with your evidence: sufficient and appropriate. Under the AICPA's AU-C Section 500 (Audit Evidence) and the PCAOB's AS 1105, sufficiency is the measure of quantity — how much evidence is needed, driven by assessed risk and the frequency of the control. Appropriateness is the measure of quality, and it has two components: relevance (does the artifact actually relate to the control being tested?) and reliability (can it be trusted, given its source and nature?). SOC 2 and SOC 3 examinations run under attestation standards (SSAE 18, AT-C section 205, and AT-C section 105 for concepts common to all attestation engagements) rather than AU-C, but the sufficient-and-appropriate framework is the same.
The single most important consequence of this framework: inquiry alone is never enough. Asking a control owner "do you review access quarterly?" is inquiry, and it cannot support a conclusion about whether the control actually operated. The auditor needs a system-generated record of the reviews, with dates and reviewers. That is why evidence management — not just evidence collection — is the operational discipline that determines audit velocity.
The audit evidence taxonomy — and how auditors rate reliability
Not all evidence is equal. Auditors classify what they see by the procedure it supports and by how reliable it is. Getting this taxonomy right is what lets you send the strongest artifact the first time instead of cycling through auditor objections.
The four test procedures evidence supports
Auditors gather evidence through four procedures, in ascending order of persuasiveness. Inquiry is asking a control owner what happens — useful for context, insufficient on its own. Observation is watching a control being performed (for example, sitting in on an access review). Inspection is examining a record or document the control produced — a ticket, a config export, a signed policy. Reperformance is the auditor independently re-executing the control (re-running an access query, recalculating a reconciliation). Design controls so they naturally produce inspection- and reperformance-grade artifacts, because those carry the most weight.
Reliability hierarchy — why system-generated reports beat screenshots
Reliability depends on a source and its nature. Evidence produced by a system under general controls, obtained directly by the auditor, is more reliable than evidence asserted by a person. A screenshot is the weakest common artifact because it is easy to stage, easy to crop, and rarely carries the population context an auditor needs. The table below shows how a CPA firm typically rates the artifacts we see most often.
| Evidence type | Procedure supported | Reliability | Example control it satisfies | Common auditor objection |
|---|---|---|---|---|
| Verbal inquiry / email confirmation | Inquiry | Low | Context only | Not sufficient to prove operating effectiveness |
| Screenshot of a UI | Inspection | Low | MFA enabled on console | No timestamp, user, or population context |
| Signed policy / procedure | Inspection | Medium | Existence of a policy (design) | Proves design, not operation |
| Ticket / approval trail | Inspection | Medium–High | Change management (CC8.1) | Missing independent approver or PR link |
| Configuration export | Inspection | High | Encryption at rest / MFA enforcement | Export not dated or not scoped to prod |
| System-generated report / immutable log | Inspection / Observation | High | Access reviews, log retention (CC7.x) | Population completeness not demonstrated |
| Auditor-reperformed query output | Reperformance | Highest | Access recert, reconciliation | Rarely disputed — auditor-obtained |
If a screenshot is genuinely unavoidable, it must capture four things in-frame to be usable: the system hostname or URL, the logged-in user identity, the full timestamp, and the scope (which environment or tenant). A cropped screenshot missing any of these will be challenged.
Point-in-time vs. period-of-time evidence
Whether one artifact suffices or you need a period-spanning sample depends entirely on the report type. A Type I report proves design as of a single date; a Type II proves operating effectiveness across a window. This is the difference between "show me the config today" and "show me every access review across the last six months."
| Report type | Evidence window | Example artifact | Sampling implication |
|---|---|---|---|
| SOC 1 Type I | Point-in-time ("as of" date) | Single config export / walkthrough | One artifact per control can suffice |
| SOC 2 Type I | Point-in-time ("as of" date) | Policy + configuration snapshot | One artifact per control can suffice |
| SOC 2 Type II | Period of time (commonly 3–12 months) | Sample of access reviews across the period | Sample must span the entire window |
| ISAE 3402 Type 2 | Period of time | Period sample of control operation | Sample must span the entire window |
The failure mode buyers and auditors call out most: a control that started mid-period, or a two-month gap in a twelve-month window. Type II evidence must be collected continuously throughout the period — gaps cannot be filled retroactively. See our companion guide on SOC 1 / SSAE 18 evidence and CUECs for the ICFR-focused equivalent.
How auditors sample evidence — sizing by control frequency
When a control operates many times over a period, the auditor cannot examine every instance. Instead they test a sample and infer operating effectiveness for the whole population. How that sample is built — and validated — is governed by AU-C Section 530 (Audit Sampling) and the AICPA Audit Guide: Audit Sampling.
Population completeness comes first
Before a single sample is pulled, the auditor must confirm the population is complete and accurate. A sample drawn from an incomplete list proves nothing, because the missing items were never eligible to be tested. This is one of the most frequent causes of expanded testing. The practical technique auditors use is reconciliation to an independent count: for an access-review population, they reconcile the review list against total headcount from the HRIS; for a change-management population, they reconcile the ticket list against merged pull requests or deployment logs. If your population source cannot be reconciled to an independent system of record, expect scope creep.
Sample size by control frequency
There is no AICPA-mandated sample size for SOC engagements. Auditors size samples from the frequency of the control and the resulting population, using the tables in the AICPA Audit Sampling Guide, which target roughly a 90% confidence level for larger populations (about 250+ items). The ladder below is an illustrative practitioner convention, not an AICPA rule; exact numbers vary materially by firm and by the confidence and expected-deviation assumptions used, so always confirm against your auditor's own methodology.
| Control frequency | Approx. population | Typical sample | Selection method | Note on deviations |
|---|---|---|---|---|
| Annual | 1 | 1 | 100% (test all) | The one instance must operate; a deviation fails the control |
| Quarterly | 4 | 2 | Haphazard | Any deviation is evaluated against the tested sample |
| Monthly | 12 | 2–4 | Haphazard | Small population: a deviation typically means the control did not operate effectively for the period |
| Weekly | ~52 | 5–9 | Haphazard / random | Deviations trigger evaluation and may expand the sample |
| Daily / many times daily | 250+ | Sized from the AICPA guide's tables (~90% confidence) | Random / statistical | Sample size and tolerable deviation set from the guide, not a fixed number |
Note the deliberate absence of hard "15–25 for daily" figures: the widely-circulated daily and multiple-times-daily numbers are not published in the common practitioner sources and should be derived from the AICPA guide's tables with a stated population, confidence level, and expected-deviation assumption — not quoted as gospel.
Handling deviations and exceptions
A deviation (or exception) is an instance where the control did not operate as designed. For a very small population where the tested sample effectively represents the control's operation — for example, a monthly control where the tested months are treated as representative — a single deviation typically leads the auditor to conclude the control did not operate effectively for the period. For larger populations, a deviation triggers evaluation: the auditor assesses whether it is isolated, expands the sample, or considers compensating controls. In all cases, document exceptions transparently. Auditors and buyers accept a disclosed exception with a corrective-action plan far more readily than a concealed one, which erodes trust and invites expanded testing. Reviewing the evidence gaps that become audit findings before fieldwork is the cheapest insurance available.
Evidence retention — how long to keep it, by framework
"How long do we keep audit evidence?" has no single answer, because retention depends on the framework and on whose records they are. There is a critical distinction between the CPA firm's workpaper retention (the auditor's own documentation) and your retention of the underlying evidence.
| Framework / party | Governing citation | Minimum retention | What must be retained |
|---|---|---|---|
| HIPAA (covered entities & business associates) | 45 CFR 164.316(b)(2)(i) | 6 years from creation or last-effective date, whichever is later | Policies, procedures, risk analyses, and records of actions/activities/assessments |
| GDPR (controllers/processors) | Art. 5(1)(e) storage limitation; Art. 5(2) accountability | No fixed number — set by policy + member-state limitation periods | Enough to demonstrate compliance; ROPA where Art. 30 applies |
| SOC 1 / SOC 2 (client evidence) | No statutory period (attestation, not statute) | Cover the full review period; retain per internal policy | All artifacts supporting each control across the period |
| CPA firm workpapers (SOC / nonissuer) | AICPA AU-C 230 / QM standards | Minimum 5 years from report release | The auditor's engagement documentation |
| ISAE 3000 / ISAE 3402 (international) | IAASB standards + firm policy | Firm-policy driven | Engagement documentation supporting the opinion |
Two facts to get right in buyer conversations. First, the "seven-year" figure people cite is the PCAOB rule (AS 1215) for issuer financial-statement audits — it does not govern SOC engagements. For a SOC report, the CPA firm's own workpapers are retained a minimum of five years under AU-C 230; your evidence of the controls is a separate matter set by your policy. Second, the frameworks: HIPAA's six-year documentation retention (45 CFR 164.316) is a hard number, whereas GDPR accountability and records of processing impose a duty to demonstrate compliance (Article 5(2)) without prescribing a single retention period — the storage-limitation principle (Article 5(1)(e)) means you keep evidence no longer than necessary, governed by policy.
One GDPR caveat worth flagging for smaller SaaS teams: the Article 30 records-of-processing (ROPA) obligation has an exemption for organizations under 250 employees — unless the processing is not occasional, is likely to result in a risk to data subjects, or includes special-category data (Art. 30(5)). Most SaaS and healthtech processing clears one of those triggers, so plan to maintain a ROPA regardless.
Multi-framework evidence: collect once, satisfy SOC 2, HIPAA, and GDPR
The controls that matter most — access management, change management, encryption, monitoring, vendor risk — recur across every major framework. The efficient program collects one artifact and reuses it against several criteria rather than running parallel evidence pulls. The map below shows the highest-value shared artifacts; for the full treatment see how to map one control to SOC 2, HIPAA, and GDPR.
| Evidence artifact | SOC 2 TSC | HIPAA cite | GDPR article |
|---|---|---|---|
| Access-list export + provisioning ticket | CC6.1, CC6.2 | 164.312(a)(1) | Art. 32 |
| Quarterly access review sign-off | CC6.2, CC6.3 | 164.308(a)(4) | Art. 32 |
| Encryption configuration export | CC6.1, CC6.7 | 164.312(a)(2)(iv), 164.312(e) | Art. 32(1)(a) |
| SIEM alert config + audit logs | CC7.1, CC7.2 | 164.312(b) | Art. 32 |
| Change ticket with independent approver | CC8.1 | 164.308(a)(8) | Art. 32 |
| Vendor risk assessment + BAA/DPA | CC9.2 | 164.308(b), 164.314 | Art. 28 |
A per-criterion evidence request specimen is worth memorizing because it is exactly what an auditor's PBC list will ask for: CC6.1 = access-list export + provisioning ticket; CC6.2 = deprovisioning report + access-review sign-off; CC6.7 = encryption config export; CC7.2 = SIEM alert configuration + a sample triggered alert; CC8.1 = change ticket linked to a PR with an independent approver; CC9.2 = completed vendor risk assessment. Where you rely on subservice organizations, decide early between the carve-out and inclusive methods, and remember that Complementary User Entity Controls (CUECs) push evidence responsibility onto your customers — you must document them so buyers know which controls they operate. See vendor and subprocessor evidence (BAAs, DPAs, subservice orgs) and, for international expectations, ISAE 3402 / international evidence expectations.
How to build an audit evidence collection process
A durable evidence program is a repeatable operating process, not a fire drill before fieldwork. These six steps take you from control register to a governed repository; for how we test and sample once you engage us, see how Auditsuisse tests and samples evidence.
- Define the control register. List every in-scope control with an owner assigned by role (roles persist; people churn), a frequency, and the criterion it satisfies.
- Map evidence to each control. Specify the exact artifact each control produces — system report, config export, ticket with approver, or reperformance output — so there is no interpretation at collection time.
- Verify population completeness. Reconcile each population to an independent system count before sampling (HRIS headcount vs. access-review list; merged PRs vs. change tickets).
- Set sample sizes by frequency. Size samples under AU-C 530 and the AICPA guide, driven by control frequency and population, with stated confidence and deviation assumptions.
- Run a pre-submission quality gate. Check every package for attribution, timestamp, scope, completeness, and reviewer independence before it reaches the auditor.
- Retain in a governed repository. Store artifacts with stable naming, version control, and permission boundaries for the retention period each framework requires.
The PBC / evidence request list and evidence calendar
The PBC list (Prepared By Client, also called the evidence request list) is the auditor's itemized ask. Build your own version proactively and tie it to an evidence calendar keyed to control frequency, so period-of-time evidence accrues continuously instead of being reconstructed at the last minute. An illustrative PBC list runs 15–30 line items, including: org chart + headcount report; onboarding/offboarding tickets sample; quarterly access-review sign-offs; MFA/SSO configuration export; encryption-at-rest and in-transit config; change-management tickets with approvals; production deployment logs; vulnerability-scan results and remediation tickets; penetration-test report; incident-response runbook + any incident tickets; backup configuration and a restore-test record; BC/DR test results; risk assessment; vendor inventory with risk ratings and BAAs/DPAs; security-awareness training completion report; policy set with version history; SIEM/logging configuration + sample alerts; and board or management security-review minutes.
The pre-submission quality gate
A short reviewer checkpoint before evidence leaves the building catches the majority of preventable deficiencies and materially improves first-pass acceptance during fieldwork.
| Check | Pass criteria | Common failure |
|---|---|---|
| Attribution (who) | Artifact shows who performed / approved the control | Anonymous export; no reviewer name |
| Timestamp (when) | Full date/time falls inside the review period | Undated screenshot; date outside window |
| Scope alignment | Artifact is from the in-scope system/environment | Evidence pulled from staging, not prod |
| Population completeness | Source reconciles to an independent count | List cannot be tied to HRIS/deployment records |
| Reviewer independence | Approver is independent of the requester where required | Self-approved change or access |
Automating evidence collection — continuous monitoring vs. manual pulls
Continuous control monitoring (CCM) platforms — Vanta, Drata, Secureframe, and similar — connect to your stack and pull evidence automatically on a schedule. They do this well for anything with a clean API surface: cloud configuration (AWS/GCP/Azure), identity providers (MFA/SSO settings, user lists), MDM (device encryption, screen lock), ticketing (change and access workflows), and HR systems (roster for onboarding/offboarding). For high-frequency technical controls, automated collection is a large improvement over screenshots because it captures attribution, timestamps, and a near-complete population by default.
But CCM has firm limits, and buyers who over-rely on it get burned. Automation cannot collect the judgment- and process-based evidence auditors still require: management review sign-offs, board or security-committee minutes, vendor risk decisions, risk-acceptance rationale, physical security, and anything requiring human attestation. And crucially, the auditor still validates the source and completeness of automated evidence — a green checkmark in a dashboard is not an auditor's opinion. The tool asserts a control passed; the auditor independently confirms the population was complete and the artifact is reliable. Treat CCM as a force multiplier for collection, not a replacement for evidence management or for the auditor's testing.
Common audit-evidence pitfalls and how to avoid them
Screenshots as the default artifact. Screenshots are the lowest-reliability evidence and rarely stand alone. Replace them with system-generated reports and config exports; when a screenshot is unavoidable, capture hostname/URL, logged-in user, full timestamp, and scope.
Incomplete populations. A sample drawn from an incomplete list is worthless. Reconcile every population to an independent system count before sampling — this single habit prevents most expanded-testing cycles.
Late or period-gap evidence. Type II evidence must accrue continuously across the window. A control that started mid-period or a multi-month gap cannot be back-filled after the fact; the exception is baked in. Run the evidence calendar from day one of the observation period.
No attribution. Evidence that does not show who acted and when cannot prove the control operated. Require attribution and timestamps in your quality gate.
Concealed exceptions. Hiding a deviation almost always backfires: auditors expand testing when they sense something is missing, and buyers lose trust. Disclose exceptions with a corrective-action plan — transparency is faster and cheaper. If you are early-stage, our guide on when to stand up an evidence program by funding stage helps sequence this without over-building.
Frequently asked questions
What is audit evidence management?
Audit evidence management is the process of collecting, organizing, retaining, and presenting the artifacts that prove a control operated as designed. Strong programs map each control to a required evidence type, verify the population is complete, sample by control frequency, and store artifacts in a governed, timestamped repository ready for auditor inspection.
What makes audit evidence sufficient and appropriate?
Under AU-C 500 and PCAOB AS 1105, evidence must be sufficient (enough of it, driven by risk and control frequency) and appropriate (relevant to the control and reliable by source). System-generated reports, logs, and reperformed results are more reliable than screenshots or inquiry; inquiry alone can never prove a control's operating effectiveness.
How are sample sizes determined in a SOC 2 audit?
The AICPA sets no fixed sample size; auditors follow AU-C 530 and the AICPA Audit Sampling Guide, sizing samples by control frequency and population, targeting roughly 90% confidence for larger populations. As an illustrative firm convention, a control tested annually might need one sample, quarterly two, monthly two to four, and weekly five or more; daily and high-frequency controls are sized from the guide's tables.
How long must you retain audit evidence?
Retention depends on the framework and on whose records they are. HIPAA requires six years from creation or last-effective date (45 CFR 164.316(b)(2)). SOC engagements have no statutory client period, but the CPA firm retains its workpapers a minimum of five years under AU-C 230. GDPR sets no fixed number—retention follows the storage-limitation principle (Article 5(1)(e)).
Are screenshots acceptable as audit evidence?
Screenshots are low-reliability evidence and rarely sufficient on their own. Auditors prefer system-generated reports, configuration exports, immutable logs, and approval trails because they carry attribution, timestamps, and population context. If a screenshot is unavoidable, capture the system hostname or URL, the logged-in user, the full timestamp, and the scope in-frame.
What is the difference between point-in-time and period-of-time evidence?
Point-in-time (Type I) evidence proves a control existed as of one date, so a single artifact can suffice. Period-of-time (Type II) evidence proves a control operated throughout a review period (commonly three to twelve months), requiring a sample of instances spanning the entire window rather than one snapshot.
Sources & further reading
- PCAOB — AS 1105, Audit Evidence (sufficient appropriate audit evidence; relevance and reliability).
- AICPA & CIMA — AU-C Section 500, Audit Evidence and AU-C Section 530, Audit Sampling; see also the AICPA Audit Guide: Audit Sampling.
- eCFR — 45 CFR 164.316(b)(2) (HIPAA six-year documentation retention) and 45 CFR 164.312(b) (audit controls).
- GDPR — Article 5 (accountability & storage limitation) and Article 30 (records of processing activities).
Need a faster path to audit-ready evidence?
Auditsuisse is a US & Swiss licensed CPA firm. We help SaaS and healthtech teams stand up an evidence program auditors can test on the first pass — complete populations, reliable artifacts, sensible retention. Explore SOC 2 evidence and the Trust Services Criteria in our compliance audit services, or book a scoping call.
Request Consultation