Back to blog SMS Quality and Operations

A2P SMS Observation Windows: How to Compare Routes Without Bias

Define cohorts, maturation periods, and closure rules to analyze A2P routes on equivalent terms, distinguish statuses from DLRs, and handle pending results cautiously.

Diagram of an A2P SMS observation window showing sent messages, received statuses, and delayed DLRs

Why a Comparison Needs a Defined Window

A rate or latency can only be interpreted if you know which messages are included in the calculation and how long you waited for their statuses. If one route is evaluated over a few hours and another over several days, delayed DLRs can change the comparison even if the initial traffic was similar.

Define the analytical window before sending or reviewing messages: the point of inclusion, the observation duration, and the rule for closing each cohort. ITU-T Recommendation E.803 explains that a limited period avoids waiting indefinitely for future events, and that events occurring after the timeout are not part of the defined calculation.

Do not confuse this window with the SMS validity period. In 3GPP TS 23.040, the validity period is a message parameter that determines how long the message remains valid at the service centre. The analytical window is an evaluation rule; it may coincide with that period, but it is not the same concept.

  • Specify when a message enters the cohort, for example, based on its send time.
  • Set how long its statuses will be observed and exactly what closure means.
  • Keep the rule so the analysis can be repeated; do not change it merely to favour a result.
Why a Comparison Needs a Defined Window

Define the Comparison Unit and Conditions

Before comparing, determine what “one route” means in the test and which message population is included. At a minimum, define the destination and operator where known, sender, traffic type—for example, OTP or transactional—and relevant sending conditions. An aggregate comparison can conceal the fact that routes received different traffic mixes.

Record variables that could differ between cohorts: time of day, volume, destination and sender composition, and sending configuration. If a factor matters to the question, keep it constant or report results separately by stratum. Stratification is a practical adaptation of quality-of-service sampling principles; the sources consulted do not provide a universal recipe specifically for A2P cohorts.

Normalizing the destination number to international format can make data preparation easier. E.164 defines the international numbering plan, but a normalized number does not establish which operator served the destination or which route the message used. Do not infer that information from the number format alone.

  • Write the cohort definition as a verifiable sentence, not just the route’s commercial name.
  • Compare equivalent populations or separate results by the variables that differ.
  • Record whether a status report was requested and what validity policy was applied; in 3GPP, the report request and validity period are optional message parameters.
Define the Comparison Unit and Conditions

Set the Start, End, and Maturation Period for Delayed DLRs

Anchor the window to an unambiguous event, usually the recorded send time. Decide whether you are evaluating a complete cohort of messages sent within an interval or each individual message over a period measured from its send time. Apply the chosen method equally to all routes.

Add a maturation period so delayed statuses have a defined opportunity to be included in the result. The cited sources do not specify a universal timeout for comparing A2P routes: the value should reflect the operational objective and analysis policy and be documented. Once the limit is reached, apply the closure rule without later adding events retroactively and selectively.

Keep message times and report times distinct. 3GPP TS 23.040 describes SMS-STATUS-REPORTs with statuses and timestamps, such as TP-SCTS and TP-DT, as well as references for correlating them with the original message. Retain these fields when available and document the meaning of any provider identifier.

  • Record the time zone and timestamp convention to avoid comparing timestamps interpreted differently.
  • Define in advance how to handle a status received after closure: exclude it from the original calculation, but retain it for an audit or clearly separate subsequent analysis.
  • If the maturation period changes between tests, present the results as evaluations under different rules, not as a direct comparison without qualification.

Keep Cohorts Comparable When Traffic Changes

A change in volume is not the only factor that can bias a result. The proportions of destinations, operators, senders, or traffic types may also vary, as may the time-of-day distribution. If one route receives more messages from one category than another, the observed difference may reflect sample composition rather than the route itself.

When the mix varies, report results by relevant stratum and retain each group’s size. If an aggregate indicator is needed, explain how it was constructed and avoid combining groups that do not represent the same target population. ITU-T E.806 describes stratified sampling for quality-of-service campaigns; applying it to A2P routes is an adapted methodological principle, not an SMS-specific requirement.

Also document operational changes that coincide with the evaluation. If routing, sender, configuration, or sending pattern changes, separate the affected periods instead of treating them as one homogeneous cohort.

  • Compare equivalent time intervals or show results by time band when timing matters.
  • Retain counts by destination, known operator, sender, and traffic type.
  • Flag configuration changes and avoid automatically attributing their effect to a single variable.

Separate Acceptance, Statuses, Latency, and Availability

Do not turn every signal into a single notion of delivery. Submission acceptance, a subsequent SMS-STATUS-REPORT, and actual receipt on a handset are different events. Acceptance alone is not equivalent to a final report, and a reported DLR should not automatically be described as independent verification that the handset received the message.

Report separately what proportion of submissions was accepted by the observed system, which statuses were received within the window, and what proportion had no observed result. Describe exactly which statuses were grouped and how. Available semantics may depend on the system that originates or transforms the report; do not assume all DLRs represent the same level of confirmation.

Associate latency with defined events—for example, the time from submission to receipt of a specific status. State which timestamp you used and, where appropriate, summarize the distribution rather than relying only on an average. Availability also requires an explicit operational definition and an observation source; do not infer it from a single status indicator.

  • Label measures according to the rule applied: acceptance, reported status, latency to a defined event, or availability.
  • State whether the data is a received and reported DLR or independent verification of handset receipt.
  • Do not treat a missing status as automatic proof of failure or a reported status as a delivery guarantee.

Incomplete Results and Small Samples

At closure, a message with no observed final status should remain identified as pending, censored, or having no observed result, according to the defined rule. Do not count it as a success or failure without evidence. State both the cohort’s total number of messages and how many still have no result; otherwise, a proportion calculated only from messages with a status can give an incomplete impression.

A small sample produces more uncertain estimates. The sources consulted do not establish a universal sufficiency threshold for comparing A2P routes. Report the sample size and, if the statistical method used allows, an appropriate measure of uncertainty. ITU-T E.802 and E.803 discuss sample size and uncertainty, but their examples are not a fixed rule for this use.

If the proportion of pending results or the observation window changed, say so alongside the metrics. A single figure without a denominator, period, or closure rule can appear precise without being comparable.

  • Show the numerator, denominator, and messages with no observed result.
  • Do not declare one route superior based on a small number of messages without noting the uncertainty.
  • Avoid treating an observed difference as a guarantee of future performance.

Minimum Record and Limits of Interpretation

A reproducible comparison requires message-level evidence and a description of the rules. Retain a correlatable identifier, the route evaluated, send time, normalized destination, and cohort variables. Add the received status, its timestamp, the observation period, the closure rule, and the number of messages without a result.

Also record conditions that affect how the result should be interpreted: whether a status report was requested, the SMS validity policy, which statuses were included in each category, and any data limitations. The message reference and timestamps described by 3GPP help with correlation; document additional provider identifiers rather than assuming they are equivalent across systems.

BulkSMSMarket describes a platform in development for discovering, comparing, buying, selling, and managing A2P SMS capacity, alongside daily internal testing of routes, destinations, and operators that observes delivery, DLR consistency, latency, availability, and sender or content behaviour. Public numerical cards are illustrative until contractual route data is connected. The operational marketplace platform, authentication, balances, billing, and live routing are not yet public; therefore, this article is an evaluation guide, not an offer of live commercial data.

  • Message identifier and route evaluated.
  • Send time, normalized destination, and cohort conditions.
  • Received status, timestamp, and classification rule.
  • Window start and end, maturation period, and closure rule.
  • Counts of accepted submissions, received statuses, and messages without a result, with their denominators.
  • Operational changes and limitations that affect interpretation.

Checklist for Repeating the Comparison

Before making a decision, verify that the routes were evaluated using equivalent criteria and that the conclusion does not depend on selectively including delayed DLRs. Changing the window may be useful for a new analysis, but it should create a new version of the result and be applied consistently.

Use the comparison as operational evidence limited to the conditions observed. Review sample composition, incomplete data, and uncertainty before changing traffic. No observation period turns past results into a guarantee of future delivery.

  • Is the comparison unit and cohort definition unambiguous?
  • Are destination, known operator, sender, traffic type, and time of day comparable or separated into strata?
  • Were the start, end, and maturation period set before interpreting the results?
  • Are delayed DLRs handled using the same rule across all routes?
  • Are acceptance, reported statuses, latency, and availability presented separately?
  • Are denominators, pending messages, and sample size shown?
  • Are the conditions, limitations, and changes recorded so the analysis can be reproduced?
FAQ

Frequently asked questions

Is the observation window the same as the SMS validity period?

No. The validity period is a 3GPP-defined message parameter referring to how long the SMS remains valid at the service centre. The observation window is an analytical rule that determines how long events are awaited when evaluating results.

Does a message accepted by the route count as delivered?

Not necessarily. Submission acceptance, receipt of a status report, and handset receipt are different events. A received DLR is a reported status; by itself, it should not be presented as independent verification of handset receipt.

What should I do with a DLR that arrives after closure?

Apply the rule defined before the analysis: do not selectively add it retroactively to the closed calculation. Retain it for an audit or separate subsequent analysis, noting that it arrived after the original window.

How many messages are needed to compare two routes?

The cited sources establish no universal threshold for this comparison. Report the sample size and uncertainty where possible, and interpret differences based on few messages cautiously.

Does a missing DLR mean the message failed?

Not automatically. At closure, classify it as pending, censored, or having no observed result according to the stated rule. Do not recode it as a success or failure without evidence.

Sources consulted

  1. 3GPP TS 23.040 V19.0.0, Release 19: Technical realization of the Short Message Service (SMS)ETSI / 3GPP
  2. ITU-T E.803 (07/2022): Quality of service parameters for supporting service aspectsInternational Telecommunication Union
  3. ITU-T E.802 (2007), Amendment 2 (06/2018): Framework and methodologies for the determination and application of QoS parametersInternational Telecommunication Union
  4. ITU-T E.806 (06/2019): Measurement campaigns, monitoring systems and sampling methodologies to monitor the quality of service in mobile networksInternational Telecommunication Union
  5. ITU-T E.164 (02/2026): The international public telecommunication numbering planInternational Telecommunication Union
  6. 3GPP specification record 23.0403GPP