Skip to main content
Change Risk Intel

SOC 2 Type II Evidence: What Auditors Actually Sample

Ben Ennis Founder, Ennis Studio · former Partner Technology Advisor, ServiceNow · 10 min read

This article is operational guidance for compliance and change owners, not legal or accounting advice. No affiliate links appear in this post.

Most teams prepare for a SOC 2 Type II audit by collecting policies and screenshots. Then fieldwork starts, the auditor asks for a population, and the gap appears. The audit does not test whether your controls exist. It tests whether they operated, every time they should have, across months, and whether you can prove the list of times they should have is complete. That is a different discipline, and it is closer to change management than most teams expect.

This piece walks through what a Type II auditor actually pulls: how the observation window works, how sample size is set, why population completeness decides more outcomes than control quality, and how an exception differs from a bad opinion. It is the companion to our SOC 2 change management controls guide, which covers the controls themselves.

Why this matters

A Type I report checks control design at a single point in time. A Type II report, defined under the same attestation standard, adds an assessment of operating effectiveness of those controls, per Wikipedia’s summary of SSAE 18 section 320. That one difference changes everything about evidence. You are no longer showing that a control was designed. You are showing that it ran, on schedule, for the entire period, and that you can produce the full record of every time it should have run. Teams that treat the audit as a document-collection exercise fail the part that actually gets tested.

The observation window is the whole game

Type II examinations are governed by AT-C section 205, Examination Engagements, one of the sections under SSAE No. 18, effective May 1, 2017, per the SSAE 18 reference. The practitioner’s objective in an examination is to obtain assurance that the subject matter is free from material misstatement and to express an opinion on whether it meets the criteria. For SOC 2 that subject matter is control operation over a period, commonly three to twelve months. There is no AICPA-defined minimum length; the engaged CPA firm approves the window, and three months is the practical floor most auditors will accept for a first Type II. The controls in scope are measured against the 2017 Trust Services Criteria, with revised points of focus, 2022, the AICPA’s control criteria for security, availability, processing integrity, confidentiality, and privacy. The practical consequence: a control that worked in month one and drifted in month four produces an exception, even if it works perfectly on the day of fieldwork.

This is why a Type II window behaves like a continuous change-control problem rather than a one-time inspection. Every in-scope change, access review, and approval during the period is a potential test item, so the discipline that keeps the audit clean is the same discipline that keeps change risk low the rest of the year: consistent execution, complete records, and a way to score which changes carry the most risk. Teams already running a structured change process, and scoring changes with something like our change risk score tool, walk into fieldwork with the population half-built, because the evidence is a byproduct of the process instead of a scramble at the end.

How auditors set sample size

Here is the fact that surprises most first-time teams: the AICPA does not require specific sample sizes, and neither does any other governing body, per KirkpatrickPrice’s sampling guidance for SOC reports. Instead, AT-C 205 directs the auditor to consider the characteristics of the control population, whether the population is homogenous, the frequency of the control’s application, and the expected deviation rate, then determine a tolerable rate of deviation and pick a sample size from there. The same guidance notes that under paragraph .32 of AT-C section 205, the auditor’s sample should be reasonably expected to represent the population across the reporting period, and random selection is one way to achieve that.

Because sample size keys off control frequency, it scales predictably even without a mandated number. The chart below shows the observed-practice ranges.

Bar chart of SOC 2 Type II sample sizes by control frequency, from one item for annual controls up to 25 to 40 for controls running many times a day, over a 6 to 12 month observation window

An annual control is tested with a single item. A quarterly control usually draws two or three, a monthly control two to five, a weekly control into the teens, and a daily or higher-frequency control from roughly 15 to 40, per practitioner ranges compiled from SOC auditor guidance. Treat these as convention, not rule. The useful takeaway for a change owner: the more often a control runs, the more evidence points the auditor will pull, so a high-frequency control with sloppy record-keeping is the fastest path to an exception.

The sampling method also determines what a single deviation means. Because the auditor picks a sample size against a tolerable rate of deviation, a control tested with 25 items tolerates a different number of misses than one tested with two before the result crosses into a reportable exception. That is not a loophole to exploit, it is a reason to be honest about control frequency when you scope the audit. Claiming a control runs daily when it really runs weekly does not help you, because the auditor will pull the larger sample and find the gaps the higher frequency implies. Scope each control at the frequency it actually operates, and make sure the evidence trail matches that claim for every occurrence in the window.

Population completeness decides the outcome

Before an auditor pulls a single sample, they confirm the population is complete. The population is the full list of every time a control should have operated during the window: every change deployed, every access review due, every access grant issued. This is where teams lose. An auditor cannot draw a representative sample from a list that is missing items, and more Type II exceptions originate in populations that could not be produced completely than in controls that genuinely failed, per Ledger Audits’ analysis of what auditors pull. A perfect approval ticket does not help if you cannot show that every in-scope event had a chance to enter the population in the first place. Completeness is what connects the sample the auditor tests to the reality the opinion covers, and it is a data-lineage problem, not a policy problem. If your change records live in one system and your deploys happen in another, reconciling the two into one provably complete population is the work.

An exception is not a failed audit

Teams panic at the word exception, but it is a narrow, technical term. An exception is an individual test deviation, recorded in Section IV of the report, the body that lists each control, the tests performed, and the results. The opinion is the auditor’s overall conclusion after weighing those deviations, and the two are not the same thing. A report can carry exceptions and still receive a clean, unqualified opinion. Schellman lists four possible SOC opinion types, with unqualified being the optimal result, meaning the assessor has no reservations; a qualified opinion means some but not all service commitments were met; and an adverse opinion means the deficiencies are both material and pervasive, per Schellman’s guide to qualified opinions. The lesson: a single documented exception with a remediation note is survivable and normal. A pattern of exceptions that shows a control did not operate as designed across the window is what moves an opinion from unqualified to qualified.

What to do about it

  1. Fix your populations before your controls. For each in-scope control, identify the source system that records every occurrence, and confirm you can export a complete, timestamped list for any date range. This is the single highest-leverage prep step.
  2. Reconcile change records against deploys. Join your change-management system to your actual deployment log so the change population is provably complete, the same discipline our risk-based patching evidence guide applies to remediation.
  3. Monitor operating effectiveness across the whole window, not at fieldwork, so drift surfaces in month two instead of during the audit.
  4. Map each control to the exact evidence artifact and the named reviewer before the window opens, and store it where you can retrieve it by date.
  5. Treat a single exception as a data point, not a crisis: document the deviation, the root cause, and the remediation, so the auditor can weigh it toward a clean opinion.
  6. Read the rest of our compliance coverage to align SOC 2, PCI, and SOX evidence into one change process instead of three.

Frequently asked questions

What does a SOC 2 Type II auditor actually sample? Evidence that each control operated across the whole window; sample size scales with control frequency, and completeness is checked first.

How long is a SOC 2 Type II observation window? Commonly three to twelve months, with no AICPA-defined minimum and three months as the practical floor.

Does the AICPA require a specific SOC 2 sample size? No. AT-C 205 sets the considerations; the published ranges are convention.

Why is population completeness so important? An incomplete population cannot yield a representative sample, and it is the leading source of Type II exceptions.

Is a SOC 2 exception the same as a failed audit? No. An exception is a Section IV test deviation; the opinion is the overall conclusion, and clean opinions can carry exceptions.

For the controls behind these criteria, see our SOC 2 change management controls guide and the wider compliance pillar.

Sources

Published September 21, 2026.