Advanced research governance ·

A Self-Storage Benchmark Is a Research Protocol, Not a Leaderboard

Define the claim, establishment, target population, frame, sample, measure, disposition, uncertainty, privacy boundary, and release gate before a benchmark result becomes persuasive.

A leaderboard can sort numbers. A benchmark has to explain why the numbers can be compared.

That distinction matters in self-storage because the industry is full of units that look similar until a study tries to count them. One operator reports at the facility level. Another submits a portfolio average. A third responds once for every brand. A fourth includes locations under construction. One source counts occupied units; another counts occupied rentable square feet. An opt-in survey attracts the operators most interested in the topic, then publishes the result as if it described the market.

The arithmetic may be flawless. The benchmark can still be structurally wrong.

A credible self-storage benchmark is not a spreadsheet with industry names down the left side. It is a versioned research protocol that defines the question, establishment, target population, frame, sample, measures, collection period, source evidence, dispositions, missingness, privacy controls, analysis, uncertainty, release rules, and correction path before the results become persuasive.

This article proposes a practical benchmark protocol for operators, associations, researchers, and technology teams. It does not publish an industry result. Every example is fictional.

Seven-gate self-storage benchmark release diagram showing claim, establishment, frame, measure, quality, privacy, and release gates with a fail path to narrow the claim, repair the protocol, or hold publication.

Start with the claim the study is allowed to make

The first design question is not “Which metrics should we collect?” It is “What claim will this study be permitted to support?”

There are at least four different products that are routinely called a benchmark:

  1. A participant snapshot describes the organizations or facilities that supplied usable data.
  2. A bounded comparison compares defined groups inside that observed sample.
  3. A population estimate attempts to infer something about a target population using a defensible frame, sample design, weights, and uncertainty.
  4. A causal study attempts to estimate what changed because of an intervention.

Those products need different evidence. An opt-in participant snapshot can be useful. It should say “among the participating facilities” and describe how participation may differ from the target population. It should not quietly become “the average self-storage facility.” A population estimate requires more than a larger response count. A causal claim requires more than a before-and-after chart.

Write the permitted claim before collecting data. If the protocol cannot support it, narrow the claim.

Define the establishment before defining the sample

The American Association for Public Opinion Research released its first dedicated Standard Definitions for Establishment Surveys in December 2025. The report treats an establishment as a single unit or location within a business or organization and emphasizes starting with a clearly defined unit of measurement. That is unusually relevant to self-storage.

A self-storage study must decide whether its unit is:

  • one physical facility;
  • one operating entity;
  • one legal entity;
  • one brand;
  • one management platform account;
  • one regional operating group;
  • one portfolio; or
  • one defined facility-period, such as a facility-month.

These are not interchangeable.

Suppose a regional manager submits one portfolio spreadsheet covering 14 facilities while 20 independent operators each submit one facility. Treating every submission as one equal record gives the portfolio one-twentieth the weight of each independent facility. Expanding the portfolio row into 14 facilities without facility-level source evidence creates 14 records that may all repeat the same aggregate value. Either choice can distort the result.

The protocol should name the unit of analysis, the reporting unit, and the observation grain separately:

  • Unit of analysis: the entity about which the study will make a statement.
  • Reporting unit: the person, system, or organizational level providing data.
  • Observation grain: the smallest record actually analyzed.

If those three are different, the study needs a documented translation rule.

Build a frozen target population and frame

The target population is the complete group the study claims to describe. The frame is the operational list from which the study can identify or select eligible units. A frame is evidence, not scenery.

For a fictional U.S. facility benchmark, the target population might be:

Active self-storage facilities operating to the public in the 50 states and District of Columbia at the frame cutoff, excluding facilities under construction, permanently closed locations, parking-only properties, portable-storage depots, and records whose facility identity cannot be resolved.

That definition still requires decisions:

  • What qualifies as active?
  • How is a facility with multiple addresses handled?
  • Is a mixed-use property eligible?
  • Are managed but not owned facilities included?
  • Which source resolves a brand-name change?
  • How are duplicate directory, operator, and vendor records reconciled?
  • What happens when a location opens after the cutoff?
  • What evidence moves a facility from unknown to eligible or ineligible?

Freeze the frame at a named timestamp and retain its version. Record births, closures, mergers, rebrands, and duplicate corrections as changes to the next frame or as governed revisions—not silent edits to the sample after outcomes are visible.

The U.S. Census Bureau's Statistical Quality Standard A3 addresses sample frames, target populations, coverage, selection probabilities, stratification, duplication, births, deaths, timeliness, and frame limitations for covered Census work. It is not a private self-storage rule. It is a useful reminder that a sample cannot be better defined than the frame from which it came.

Separate a census, a probability sample, and an opt-in sample

The study must say how units entered the data.

Attempted census

An attempted census invites or collects every known eligible unit in the frame. It is not automatically complete. Missing facilities, nonresponse, invalid records, and unresolved eligibility can still create coverage and nonresponse error.

Probability sample

A probability sample gives every eligible unit a known, non-zero selection probability. That permits design-based inference when the frame, selection, weights, response, and variance estimation are handled correctly. Stratification may improve coverage or precision, but it must be reflected in weights and uncertainty.

Non-probability or opt-in sample

An opt-in sample contains units that volunteered, were recruited from a non-probability source, or otherwise lack known selection probabilities. It can support a participant description and carefully bounded comparisons. It does not become representative because it contains recognizable operators or a large number of facilities.

AAPOR's current Transparency Initiative asks researchers to disclose whether a sample is probability-based, non-probability, AI-generated, or combined; describe the frame and recruitment; identify uncovered segments; disclose incentives; and avoid measures of precision that the design cannot support. The initiative explicitly says it promotes methodological disclosure and does not judge the quality or rigor of the method. Citing its disclosure elements does not make this proposed protocol an AAPOR standard, and it does not imply Jared or any organization is a Transparency Initiative member.

Keep a disposition for every sampled establishment

A final dataset hides the paths that did not produce a usable row. A benchmark protocol keeps them visible.

For every sampled or invited establishment, record a final disposition such as:

  • eligible complete;
  • eligible sufficient partial;
  • eligible insufficient partial;
  • eligible refusal;
  • eligible noncontact;
  • ineligible;
  • duplicate;
  • closed before the reference period;
  • unresolved eligibility;
  • source unavailable; or
  • excluded under a predeclared rule.

These are proposed operating labels for the companion worksheet, not copied AAPOR codes.

Consider a fictional sample of 240 facility records. After review, seven are duplicates, 18 are ineligible, 11 are insufficient partials, 142 are complete, and 62 are eligible nonresponses. Calling the response rate “142 divided by 240” ignores how the protocol treats duplicates, ineligibility, partials, and unknown eligibility. The correct rate depends on the stated disposition framework and formula.

AAPOR's December 2025 establishment-survey report is particularly useful here because it focuses on final dispositions for business and organizational locations. Its definitions are professional survey-research guidance. They do not validate a self-storage study, set an industry threshold, or eliminate the need to publish the actual formula.

Freeze the metric dictionary before looking at the result

The metric-definition card comes before the benchmark table.

For every reported measure, freeze:

  • metric name and plain-language purpose;
  • unit of analysis and observation grain;
  • source system and source owner;
  • reference period, cutoff, and timezone;
  • inclusion and exclusion rules;
  • numerator and denominator;
  • aggregation method;
  • weighting rule;
  • missing, invalid, duplicate, and late-record treatment;
  • correction and restatement policy;
  • transformation version; and
  • known limitations.

Self-storage-specific definitions deserve particular care.

Occupancy

State whether the measure is unit occupancy, rentable-square-foot occupancy, economic occupancy, or another defined construct. Name how offline, damaged, model, employee, auction, overlocked, and unrentable units enter the population. Do not compare a snapshot at one cutoff with a monthly average unless the study defines that transformation.

Delinquency

Name whether the unit is an account, agreement, tenant, facility, balance, or dollar. Define the age threshold, charge treatment, payment posting cutoff, credits, write-offs, reversals, and denominator.

Leads and conversion

Define whether one person with a call, form, and chat is three events, one contact, or one inquiry episode. State qualification, deduplication, attribution window, cancellation handling, and the exact conversion event.

Rate and revenue measures

Name taxes, insurance, fees, discounts, promotions, concessions, unit mix, tenure, and whether the reported value is offered, contracted, billed, collected, or reconciled.

AI-assisted measures

Separate the model output from the human decision and the operating outcome. Publish the test population, inclusion rule, reference answer or adjudication method, model and prompt version, human-review rule, uncertainty, and limitations. Synthetic or AI-generated responses must not be presented as human participant data.

Capture source evidence and transformation lineage

A benchmark row should be reconstructable from the governed source.

For each participating facility or portfolio contribution, retain:

  • stable facility and reporting-unit identifiers;
  • source name and source owner;
  • extraction time;
  • source-period boundaries;
  • raw record count;
  • schema and export version;
  • mapping and join keys;
  • validation checks;
  • corrections and exclusions;
  • transformation code or formula version;
  • analyst and reviewer; and
  • a safe evidence reference or checksum.

Do not place tenant names, payment details, gate logs, credentials, or customer-identifying records into the public reproduction package. Reproducibility does not require publishing sensitive raw data. It requires publishing enough method, code, schema, synthetic examples, aggregate inputs, and verification evidence for another qualified reviewer to understand and test the process.

The GAO's 2019 Assessing Data Reliability guide describes accuracy, completeness, and applicability for the intended purpose and uses a flexible, risk-based assessment. That is audit guidance, not a private benchmark certification. The useful operating lesson is narrower: “reliable” is incomplete without “reliable enough for which purpose, based on which review?”

Measure coverage, missingness, and nonresponse

Publication should show the data that did not make it into the headline.

At minimum, disclose:

  • target-population definition;
  • frame size and cutoff;
  • sampled or invited units;
  • duplicates removed;
  • known eligible and ineligible units;
  • unknown-eligibility units;
  • complete and partial contributions;
  • item-level missingness for key metrics;
  • facility and portfolio coverage by relevant strata;
  • records corrected, imputed, or excluded;
  • unmatched records after linkage;
  • weights and trimming, if used; and
  • sensitivity analysis where consequential.

The Census Bureau's Standard D3 and Standard F2 provide extensive examples of documenting coverage, unit and item nonresponse, allocation or edit changes, invalid data, record-linkage losses, methods, anomalies, corrective actions, and limitations for Census products. A private study should not copy Census thresholds or imply federal compliance. It can adopt the discipline of making data-quality indicators visible beside the result.

Treat privacy as a study property, not a final redaction step

Benchmark data can expose more than the released table suggests. A small operator, unusual market, rare unit mix, distinctive event, or linked public record may make an apparently anonymous row identifiable.

The protocol should define:

  • data minimization at collection;
  • participant notice and permitted uses;
  • access by role;
  • separation of identity keys from analysis data;
  • retention and deletion rules;
  • encryption and secure transfer;
  • aggregation and minimum-cell release rules;
  • suppression and complementary suppression;
  • review of free text;
  • re-identification testing proportional to risk;
  • vendor and model data-use boundaries; and
  • incident, withdrawal, correction, and revocation handling.

NIST's Privacy Framework 1.0 is the current final framework on the official program page; version 1.1 remains an initial public draft with a final release described as coming soon. The framework is voluntary and cross-sectoral. NISTIR 8053 explains that de-identification can reduce privacy risk while also warning that some de-identified data can be re-identified. Neither source makes a dataset safe by declaration or substitutes for applicable legal, contractual, privacy, and security review.

Predeclare the analysis and uncertainty

An analysis plan written after the results are known is an explanation. An analysis plan frozen before results are visible is a control.

Predeclare:

  • primary and secondary measures;
  • comparison groups;
  • minimum usable sample rules;
  • weighting and calibration;
  • outlier and influence treatment;
  • imputation;
  • subgroup publication rules;
  • multiplicity handling where relevant;
  • uncertainty estimation;
  • sensitivity analyses;
  • correction policy; and
  • conditions that will stop publication.

If the sample is non-probability, say what uncertainty statements are and are not supportable. Do not attach a conventional margin of sampling error to an opt-in sample merely because software can calculate one. If a model-based adjustment is used, publish the model, assumptions, diagnostics, and limitations.

Separate descriptive differences from causal claims. A higher-performing group may differ in market, facility age, unit mix, operator scale, capital, systems, customer mix, or data quality. The benchmark can describe the observed difference. It cannot name the cause without a design that supports causal inference.

Govern AI across the research lifecycle

AI can help draft questions, classify free text, map schemas, detect anomalies, write code, summarize limitations, or produce synthetic test data. Each use changes the evidence path.

The AAPOR report Responsible AI Integration in Survey Research, released in May 2026, proposes a disclosure framework across the survey lifecycle and distinguishes AI-supported research from human participant data. It was commissioned, reviewed, and accepted by AAPOR's Executive Council as a service to the profession, while the report states that its opinions are those of the authors. That limitation matters.

For each AI use, record:

  • purpose and stage;
  • model and version;
  • provider and data-use terms;
  • input data classification;
  • prompt or instruction version;
  • human reviewer;
  • evaluation method;
  • known error modes;
  • changes accepted or rejected;
  • reproducibility limitation; and
  • public disclosure.

Never let synthetic responses increase the human response count. Never let model-generated categories become a validated codebook without human evaluation. Never let an AI summary replace the underlying disposition, missingness, or limitation table.

Publish a reproduction package, not just a PDF

A useful release package can contain:

  1. the frozen protocol and version history;
  2. sponsor, funding, investigator, conflict, and AI-use disclosures;
  3. target-population and frame description;
  4. sampling and recruitment method;
  5. instrument and question wording;
  6. metric dictionary;
  7. disposition and data-quality table;
  8. transformation and analysis code;
  9. synthetic or public-safe example data;
  10. weights and variance method, when applicable;
  11. privacy and disclosure rules;
  12. results with uncertainty and limitations;
  13. correction and withdrawal policy; and
  14. persistent version and citation information.

The AAPOR Transparency Initiative calls for core methodological disclosures when findings are released and additional details upon request. Again, disclosure is not a quality badge. It makes the study inspectable.

A seven-gate release decision

Use the companion 40-control benchmark protocol register and the seven-gate release diagram. The register moves from beginner claim and metric decisions through architect-level sampling, lineage, uncertainty, privacy, AI, reproduction, and correction controls. Every populated row is fictional. Stop the release unless all seven gates pass:

  1. Claim gate: The intended language matches the design.
  2. Establishment gate: Unit of analysis, reporting unit, and observation grain are explicit.
  3. Frame gate: Target population, frame version, coverage, eligibility, and duplicate rules are frozen.
  4. Measure gate: Definitions, periods, sources, formulas, and corrections are reproducible.
  5. Quality gate: Dispositions, missingness, linkage losses, validation, weights, and uncertainty are visible.
  6. Privacy gate: Uses, permissions, minimization, access, aggregation, and re-identification risks are reviewed.
  7. Release gate: Methods, limitations, conflicts, AI use, version, correction path, and reproduction package ship with the result.

If a gate fails, the study can still release a narrower product. A population benchmark may become a participant snapshot. A causal claim may become a descriptive comparison. A public microdata file may become a synthetic example plus controlled-access review. Narrowing the claim is a sign of research control, not failure.

The standard I would expect before trusting a self-storage benchmark

I would not ask whether the leaderboard looks plausible. I would ask whether another qualified reviewer can reconstruct the population, sample, metric, transformation, analysis, uncertainty, privacy boundary, and release decision without a private explanation from the sponsor.

If the answer is no, the artifact may still be a useful market conversation. It is not yet a reproducible benchmark.

That distinction is worth protecting. Operators will make staffing, pricing, marketing, capital, automation, and product decisions from industry numbers. The obligation is not to make the chart look certain. It is to make the evidence boundary visible enough that the reader knows what the number can—and cannot—support.

Sources and limitations

Observed August 22, 2026:

These sources inform individual controls. AAPOR material addresses survey and public-opinion research. Census and OMB standards govern specified federal work. GAO material is audit guidance. NIST material is voluntary privacy and risk guidance. None defines a self-storage benchmark, certifies this proposed protocol, validates a future study, establishes legal compliance, or endorses Jared Mastroianni, modSTORAGE, Facily.ai, or Facily OS.

Disclosure

I am Chief Operating Officer of modSTORAGE and CEO and Co-Founder of Facily.ai. This is an authored proposed research protocol. It is not a company study, released benchmark, customer case, product capability, performance result, legal or statistical opinion, professional-association standard, certification, or independent research finding.

About the author

Jared Mastroianni

Chief Operating Officer of modSTORAGE and CEO and Co-Founder of Facily.ai. Jared writes from the intersection of self-storage operations, accountable artificial intelligence, and operator-shaped software.