ARS and ARDs in Action

R in Pharma European Summit 2026

Introductions

Orla Doyle 🇮🇪

  • Role: Executive Director
    Head of Scientific Product Engineering
    AQS
  • Organisation: Novartis
  • Focus area: Develops workflows and products for clinical trial reporting
  • Based in: Dublin, Ireland

Gregory Chen 🇨🇭

  • Role: Principal Statistician (Director)
    BARDS HTA Statistics
  • Organisation: MSD
  • Focus area: Integrated evidence planning and synthesis via statistical techniques for HTA/value proposition/market access
  • Based in: Basel, Switzerland

Demystifying ARS

How we got here

ARS?

1

New idea
Let’s use ARS

A wall of interconnected notes

2

ARS is complex

Building a custom approach

3

Build our own

Regret after building our own

4

Despair

CDISC Analysis Results Standard logo

5

ARS v1.0 adoption

A familiar friend in clinical trial reporting - the adverse events table

Treatment-emergent adverse events by primary system organ class and preferred term

Adverse Event Treatment A
(N=xx)
Treatment B
(N=xx)
Total
(N=xx)
Number of subjects with at least one event xx (xx.x) xx (xx.x) xx (xx.x)
Cardiac disorders xx (xx.x) xx (xx.x) xx (xx.x)
 Sudden death xx (xx.x) xx (xx.x) xx (xx.x)
General disorders and administration site conditions xx (xx.x) xx (xx.x) xx (xx.x)
 Sudden death xx (xx.x) xx (xx.x) xx (xx.x)
 Postoperative fever xx (xx.x) xx (xx.x) xx (xx.x)

Generating the Table: Parameters Are Embedded

Typical implementation — key parameters hardcoded directly in the script:

library(tidyverse)

# Safety population

adsl_safe <- adsl %>% 
  filter(SAFFL == "Y")              # <- hardcoded

n_trt <- count(adsl_safe, TRTA)    # <- hardcoded

# Treatment-emergent AEs
adae_te <- adae %>%
  filter(
    USUBJID %in% adsl_safe$USUBJID,
    TRTEMFL == "Y"                  # <- hardcoded
  )

# Count by System Organ Class and Preferred Term
result <- adae_te %>%
  group_by(AESOC, AEDECOD, TRTA) %>%  # <- hardcoded
  summarise(n = n_distinct(USUBJID), .groups = "drop") %>%
  left_join(n_trt, by = "TRTA") %>%
  mutate(
    pct  = round(n.x / n.y * 100, 1),
    cell = sprintf("%d (%.1f)", n.x, pct)
  )

Every filter, variable name, and hierarchy is buried in the code.
Changing the population means hunting through the script.

Separating Parameters from Code

Pull parameters into a YAML metadata file — the code becomes generic:

analysis_config.yaml

population:
  dataset:  ADSL
  variable: SAFFL
  value:    "Y"

events:
  dataset:  ADAE
  variable: TRTEMFL
  value:    "Y"

grouping:
  variable: TRTA

subject_id: USUBJID

hierarchy:
  - AESOC
  - AEDECOD

R code — driven by metadata

library(tidyverse)
library(yaml)

cfg <- read_yaml("analysis_config.yaml")

# Population — from metadata
adsl_safe <- adsl %>%
  filter(.data[[cfg$population$variable]] ==
           cfg$population$value)
n_trt <- count(adsl_safe,
               .data[[cfg$grouping$variable]])

# Events — from metadata
adae_te <- adae %>%
  filter(
    USUBJID %in% adsl_safe$USUBJID,
    .data[[cfg$events$variable]] ==
      cfg$events$value
  )

# Hierarchy — from metadata
result <- adae_te %>%
  group_by(across(all_of(
    c(cfg$hierarchy, cfg$grouping$variable)
  ))) %>%
  summarise(
    n = n_distinct(.data[[cfg$subject_id]]),
    .groups = "drop"
  )

Metadata Quickly Grows

One AE table means many analyses. The YAML expands rapidly:

reporting_event:
  id: "RE-AE-001"
  label: "AE Summary Table"

  analysis_sets:
    - id: AS001
      label: "Safety Analysis Set"
      condition:
        dataset:    ADSL
        variable:   SAFFL
        comparator: EQ
        value:      "Y"

  data_subsets:
    - id: DS001
      label: "Treatment-Emergent AEs"
      condition:
        variable: TRTEMFL
        value:    "Y"
    - id: DS002
      label: "Serious TEAEs"
      conditions:
        - {variable: TRTEMFL, value: "Y"}
        - {variable: AESER,   value: "Y"}
    - id: DS003
      label: "TEAEs leading to D/C"
      conditions:
        - {variable: TRTEMFL, value: "Y"}
        - {variable: AEACN,   value: "WITHDRAWN"}
  groupings:
    - id: GF001
      groupingVariable: TRTA
      dataDriven: false
      groups:
        - {id: G1, label: "Treatment A"}
        - {id: G2, label: "Treatment B"}
        - {id: G3, label: "Total"}

  methods:
    - id: M001
      label: "n and percent"
      operations:
        - {id: OP1, label: "n",
           function: COUNT_DISTINCT}
        - {id: OP2, label: "%",
           function: PERCENT_BY_SUBSET}

  analyses:
    - {id: AN001, label: ">=1 TEAE",
       dataset:     ADAE,
       analysisSet: AS001,
       dataSubset:  DS001,
       grouping:    GF001,
       method:      M001}
    - {id: AN002, label: "Serious TEAE",
       dataset:     ADAE,
       analysisSet: AS001,
       dataSubset:  DS002,
       grouping:    GF001,
       method:      M001}
    # ... 30+ more analyses per table
    # ... 100s of tables per study

Without a schema, this is unvalidated.
Without a standard, every team invents their own format.

Analysis Results Standard

ARS v1.0: Background

What it defines

  • A formal metadata model for clinical trial analysis results
  • Connects ADaM datasets (inputs) to results (outputs)
  • Covers concetps like populations, filters, groupings, methods, operations

How it helps

  • Unambiguous — no programmer interpretation required
  • Reproducible — any compatible tool gives the same result
  • Auditable — every result traces to a specification
  • Automatable — machines can execute the metadata directly

The ARS Stack

Layer Purpose
LinkML Tooling use to house the blueprint that defines and validates the ARS structure
ARS (before execution) A machine-readable analysis specification
ARS (after execution) The same specification, now enriched with results
ARD The analysis results represented as structured result records that can be extracted from the completed ARS

The Wider ARS Landscape

The AE example is one slice of a broader results model: from study context and data inputs, through analysis definitions, to governed results and provenance.

Study and data context

  • Study and reporting event
  • Dataset and variable references
  • Controlled terminology
  • Analysis populations and subsets

Analysis definition

  • Analysis
  • AnalysisSet
  • DataSubset
  • GroupingFactor
  • Analysis method and operations

Results and hand-off

  • Result groups and values
  • Counts, percentages, summaries
  • Structured output records
  • Hand-off to tables, listings, and figures

Traceability and governance

  • Specification-to-result links
  • Versioning and provenance
  • Validation and conformance
  • Reproducible execution

The automation gap: ARS can make the result computable and traceable, but the final rendering and presentation metadata may live outside the model. Turning an ARD into a compliant table, listing, or figure still needs another governed layer.

This deck zooms in on the centre: how a familiar AE table becomes a structured, executable analysis definition.

Demystifying ARS, ARD, and LinkML

ARS = Recipe
The specification that describes what to analyze and how to compute it

ARD = Cooked Meal
The actual results produced by executing the ARS specification

LinkML = Recipe Template
Tooling to check the schema defining the structure and rules for writing valid recipes

Key Insight: ARS starts as a specification document. After execution, it can also contain the resulting ARD — the specification and its results travel together.

Worked Example: From AE Spec to Results

ARS Specification (Recipe)

analysis:
  id: AN001
  label: "Subjects with >=1 TEAE (%)"
  dataset: ADAE
  analysisVariable: USUBJID

analysisSet:
  condition: ADSL.SAFFL == "Y"

dataSubset:
  condition: ADAE.TRTEMFL == "Y"

grouping:
  variable: TRTA

method:
  operation: Operation_rollup_catvar_summ_grp_pct
  numerator: COUNT_DISTINCT(USUBJID) in dataSubset by TRTA
  denominator: COUNT_DISTINCT(USUBJID) in analysisSet by TRTA

↓ Execute

ARD (Cooked Meal)

analysis_id analysisSet dataSubset group1_variable group1_value operation_id raw_value formatted_value
AN001 ADSL.SAFFL EQ Y ADAE.TRTEMFL EQ Y TRTA Treatment A Operation_rollup_catvar_summ_grp_pct 0.25 25.0
AN001 ADSL.SAFFL EQ Y ADAE.TRTEMFL EQ Y TRTA Treatment B Operation_rollup_catvar_summ_grp_pct 0.33 33.0

✓ Results

LinkML ensures both the specification and results follow the same schema — any tool that understands ARS can read and execute this.

ARS Model: Core classes for ARD generation

%%{init: {'theme':'default', 'flowchart': {'nodeSpacing': 24, 'rankSpacing': 32}, 'themeVariables': { 'fontSize':'12px', 'fontFamily':'arial'}}}%%
classDiagram
  direction LR

    class ReportingEvent {
    id
    label
    }
    class Analysis {
    id
    dataset
    analysisVariable
    }
    class AnalysisSet {
    condition
    }
    class DataSubset {
    condition
    }
    class GroupingFactor {
    groupingVariable
    }
    class AnalysisMethod {
    label
    }
    class Operation {
    label
    resultPattern
    }

    ReportingEvent --> Analysis : contains
    Analysis --> AnalysisSet : analysisSet
    Analysis --> DataSubset : dataSubset
    Analysis --> GroupingFactor : orderedGroupings
    Analysis --> AnalysisMethod : method
    AnalysisMethod --> Operation : operations

AnalysisSet = who (population)

DataSubset = which records (event filter)

GroupingFactor = how to split (columns)

Operation = what to compute

Mapping ARS Back to Our Table

Every element of the AE table maps to an ARS class:

Adverse Event
dataset: ADAE  AnalysisSet: SAFFL=“Y”  DataSubset: TRTEMFL=“Y”
Treatment A (N=xx)
GroupingFactor: TRTA
Treatment B (N=xx)
GroupingFactor: TRTA
Total (N=xx)
GroupingFactor: TRTA
Subjects with ≥1 TEAE, n (%) COUNT xx  PCT (xx.x) COUNT xx  PCT (xx.x) COUNT xx  PCT (xx.x)
Cardiac disorders xx (xx.x) xx (xx.x) xx (xx.x)
 Palpitations  analysisVariable: USUBJID xx (xx.x) xx (xx.x) xx (xx.x)
General disorders and administration site conditions xx (xx.x) xx (xx.x) xx (xx.x)
 Fatigue  analysisVariable: USUBJID xx (xx.x) xx (xx.x) xx (xx.x)
  • dataset ADaM source
  • AnalysisSet population
  • DataSubset row filter
  • GroupingFactor columns
  • analysisVariable unit counted
  • Operation computation

Demo: ARS-Driven Automation for CSRs

ARS and ARDs in Action: Market Access

Same Recipe, New Kitchen, New Diners

ARS: the recipe populations, subsets, groupings, methods, operations as data

ARD: the meal the result records an executed recipe yields

LinkML: the template what a valid recipe looks like

The CSR kitchen cooks for one diner: the regulator

One base ingredient, the trial data (SDTM, ADaM). One set menu, fixed by the sponsor in the SAP. One meal: the CSR tables and the integrated summaries (ISS, ISE).

Market access serves a table of HTA assessors

Each with their own tastes and allergies: a PICO (population, intervention, comparator, outcome), a template, and evidence they will not accept. Same base ingredient, plus ingredients from outside the trial, often not grown by the sponsor: published aggregates from a literature review, real-world data, registries, expert elicitation. The recipe has to adapt, not just re-run.

In the first published Joint Clinical Assessment, the HTA assessors could not reproduce most of the naive safety comparisons. The error was in an ingredient: event rates extracted from the comparator’s publication had been switched between two outcomes. It was caught because an HTA assessor tried to recompute the numbers from the dossier and could not. [1]

[1] Joint Clinical Assessment Report, Tovorafenib, v1.0, endorsed by the HTA Coordination Group 30 Apr 2026, published 9 Jun 2026, Appendix D.2, request 4 (p. 138): “The assessors were not able to reproduce the effect estimates shown for the naïve comparisons for most of the safety outcomes”; developer’s response: “the event rates for these two outcomes in the Bouffet 2023 study had been switched around in the analyses”. health.ec.europa.eu/publications/joint-clinical-assessment-report-tovorafenib-ojemda_en

Two Asks of the Same Base Ingredient

%%{init: {'theme':'default', 'themeVariables': {'fontSize':'16px', 'fontFamily':'Source Sans Pro, arial'}}}%%
flowchart LR
  T["<b>Base ingredient</b><br/>trial data, one or more studies<br/>SDTM, ADaM"]
  X["<b>Added ingredients</b><br/>literature aggregates, RWD,<br/>registries, expert elicitation"]
  subgraph HTA["HTA track"]
    direction LR
    H1["<b>EU JCA</b><br/>relative clinical effectiveness<br/>and safety per PICO, with its<br/>certainty; no value judgement"] --> H2["<b>National HTA</b><br/>added benefit, cost-<br/>effectiveness, budget impact"] --> H3(["<b>Price, reimbursement,</b><br/><b>patient access</b>"])
  end
  subgraph REG["Regulatory track"]
    direction LR
    R1["<b>EMA</b><br/>quality, safety, efficacy;<br/>benefit–risk balance"] --> R2(["<b>Marketing</b><br/><b>authorisation</b>"])
  end
  T ~~~ X
  T --> R1
  X -. very occasionally .-> R1
  X --> H1
  T --> H1
  classDef reg fill:#eff6ff,stroke:#1e3a8a,color:#1a1a1a;
  classDef hta fill:#fef5e7,stroke:#e67e22,color:#1a1a1a;
  classDef ev fill:#f0fdf4,stroke:#16a34a,color:#1a1a1a;
  classDef ext fill:#faf5ff,stroke:#7c3aed,color:#1a1a1a;
  class R1,R2 reg;
  class H1,H2,H3 hta;
  class T ev;
  class X ext;

Regulator asks HTA body asks
The question Do benefits outweigh risks for the indication? How much better than what we use today, for whom, and is it worth the price?
Comparator, population The comparator justified at design; the trial population The local standard of care; the population the local pathway treats, often a subpopulation
Ingredients Trial data, one or more studies; RWE the exception Trial data, often several data cuts, re-analysed per template; plus SLR aggregates, real-world data, expert elicitation, costs and utilities: assembled by the sponsor, often not owned
Output One body, one authorisation, one label Many bodies, one dossier each: own populations, added-benefit category, cost per QALY, price and reimbursement

Regulatory track: once, EU-wide. HTA track: JCA in parallel with the EMA procedure, then national HTA, pricing and reimbursement country by country (EU mean 597 days from authorisation to availability, EFPIA W.A.I.T. 2025). Reg. (EU) 2021/2282, Recital 4, Art. 1(2), 2(6); Eichler 2010.

The EU Joint Clinical Assessment in One Picture

%%{init: {'theme':'default', 'themeVariables': {'fontSize':'13px', 'fontFamily':'Source Sans Pro, arial'}}}%%
timeline
    title Regulation (EU) 2021/2282 roll-out
    Jan 2025 : New active substances for cancer
             : Advanced therapy medicinal products
    Jan 2028 : Orphan medicines
    Jan 2030 : All other centrally authorised medicines

How it runs [1]

  • Member States’ PICOs are consolidated into one scope with several PICOs; one JCA dossier is filed in parallel with the EMA application
  • A published report on relative effectiveness and safety, no value judgement, no price; national bodies “give due consideration”, then decide benefit, price and reimbursement
  • Commission tracker, extract of 10 Sep 2026: 23 started since January 2025, 4 published, 16 running, 3 discontinued (two for failing the Article 9 dossier requirements); published reports carried 1, 7, 8 and 12 PICOs [2]

Key clinical evidence in the dossier [3]

  • Per PICO: efficacy, safety, all subgroups with interaction tests, AEs by SOC and PT; each analysis flagged as prespecified or not, multiplicity-controlled or not, and who requested it
  • Where no head-to-head trial answers a PICO: indirect comparisons (Bucher, NMA, MAIC, STC) with their input data, model choices and convergence checks
  • CSRs, protocols and analysis plans as appendices; no individual patient data; code only “if the analyses and corresponding calculations cannot be described by a specific standard method”

[1] Regulation (EU) 2021/2282, Art. 7(2), 8, 9(1), 11, 13. [2] European Commission, JCA tracker (ongoing, completed, discontinued), data extracted 10 Sep 2026; published JCA reports: lurbinectedin (1 PICO), tarlatamab (7), tovorafenib (8), onasemnogene abeparvovec (12 in the final scope). [3] HTACG, guidance on the JCA dossier template for medicinal products (adopted 28 Nov 2024), App. A.2, A.3, B.2, D.3, D.5; guidance on multiplicity, subgroup and sensitivity analyses (10 Jun 2024).

What HTA Analytics Look Like in Practice

Volume

  • German benefit dossiers grew from about 750 to 3,500 pages on average after the 2020 template change; outliers reach 20,000 [1]
  • Mandatory subgroups (sex, age, severity, region, stratification factors) for every endpoint [2]
  • Safety by SOC and PT with RR, CI and p, repeated per subgroup: one safety annex ran to 926 pages [3]
  • IQWiG sets aside 33% of AE analyses and 73% of subgroup categories submitted [4]

Ingredients that are not trial data: a meal rebuilt from its photograph

  • Comparators the trial never had: literature aggregates, RWD and registry controls, expert elicitation [5]
  • Survival curves digitised from a PDF and turned back into patients who never existed, then fed to NMA or population adjustment [6]
  • Three of the first four published JCAs rest on indirect comparisons; one on an effective sample size below six [7]
  • Thresholds are defaults, not a shield: 10 patients per subgroup, 10 events in one, interaction p < 0.05, incidence ≥ 5%. Filtered at 5%, tarlatamab’s JCA has no SOC breakdown of AE deaths [2,8,11]

Outputs that are not tables

  • Extrapolation parameters, transition probabilities and utilities feed a national cost-effectiveness model, outside the JCA’s clinical scope; NICE requires it fully executable, and its HTA assessors re-analyse it in 93 of 100 appraisals [9,10]
  • Countries file on different data cuts. Nobody requires an index of everything produced; the company needs one to know the totality of its own evidence
  • A dossier is read by HTA assessors who see at most patient listings, never analysis datasets they could re-run

[1] Schweitzer et al., Eur J Health Econ 2024;25(5): Modules 1–4 from 750 to 3,500 pages on average, “in individual cases 20,000 to even 40,000 pages”. [2] G-BA Module 4 dossier template, version 18 Nov 2025, §4.2.5.5, §4.3.1.3. [3] G-BA, empagliflozin Modul 4A Anhang 4-G, 28 Mar 2022, 926 pages; nintedanib Anhang 4-G, Feb 2025, 349 pages. [4] Schweitzer et al. 2024 (vfa, Dec 2022: 77% of submitted analyses not considered). [5] HTACG methodological guideline for quantitative evidence synthesis (8 Mar 2024); NICE DSU TSD 26 (expert elicitation). [6] Guyot et al., BMC Med Res Methodol 2012;12:9; HTACG practical guideline (8 Mar 2024). [7] Published JCA reports, Jun–Sep 2026 (tovorafenib ESS 5.82). [8] HTACG JCA dossier guidance (28 Nov 2024), App. A.2; IQWiG General Methods 8.0 §10.3.9. [9] NICE technology appraisal manual PMG36 §5.5.15. [10] Carroll et al., Value Health 2017;20(6): 100 single technology appraisals, exploratory analyses of the company’s model. [11] Tarlatamab JCA report, Table 98 note a.

The Analysis Space Is Combinatorial

Analysis families

  • Descriptive: baseline, disposition, drug use before, during, after
  • Efficacy, PRO, quality of life
  • Safety: overall, serious, grade ≥3, by SOC and PT
  • Subgroups, interaction tests, sensitivity
  • Model inputs: transition probabilities, PRO-mapped utilities

SAP and each HTA template

×

Populations

  • Label population
  • Subpopulations along the local pathway: line, biomarker, prior therapy, each with its own comparator

Member States’ PICOs, national pathways

×

Comparators and their ingredients

  • Trial control arm
  • Standard of care per country: indirect comparison on literature aggregates, digitised curves
  • External controls: RWD, registries
  • Expert elicitation where data are absent

the HTA assessor, not the sponsor

×

Data cut-offs

  • Regulatory cut
  • Later cuts for each region’s timing
  • Different countries file on different cuts

submission calendars

×

Templates

  • EU JCA dossier
  • G-BA Module 4
  • NICE, HAS, AIFA, …
  • Own effect measures, thresholds, layouts

each HTA body

= thousands of result records, each one analysis serving a PICO: population, intervention, comparator, outcome. ARS names the same ideas differently: analysis set, grouping, data subset, method. Same recipe structure, different vocabulary: aligning it, and extending it where HTA has words ARS does not, is the next two slides.

What ARS Names, and What HTA Would Add

ARS class In the earlier AE table In the HTA setting What to record, for whom [1,2]
ReportingEvent The CSR One dossier per body, data cut and template version Which body, cut and template: to find the right dossier, and the delta between two fits
AnalysisSet, DataSubset SAFFL = “Y”, TRTEMFL = “Y” The PICO population, often a subpopulation of the label; event subsets unchanged Which PICO the population answers: so the HTA assessor can find the analysis they asked for fits
GroupingFactor TRTA; SOC and PT, data-driven Treatment vs the local comparator; effect-modifier subgroups The comparator’s identity and where its data came from: trial arm, publication, registry, elicitation stretch
AnalysisMethod, Operation COUNT, PERCENT RR with CI and p; weighted HR after matching; network meta-analysis The assumptions behind the estimate: matched effect modifiers, weights and effective sample size, priors, convergence. The JCA checklist asks for exactly these stretch
Analysis.reason, purpose SPECIFIED IN SAP; PRIMARY OUTCOME MEASURE Requested by a Member State; post hoc for PICO 5; requested by an HTA assessor Prespecified or not, multiplicity-controlled or not, and who asked: required per analysis by the JCA guidance. ARS has four reason terms, none of them HTA stretch
OperationResult, ResultGroup n (%) per cell RR, CI, p per PICO × subgroup × cut Identifiers tying each number to its PICO, cut and template: for navigation, and for a model to cite a record instead of a page fits
Programming code, context The R script Code per analysis, as JCA Appendix D.3 asks The code and its environment: reproducibility for the sponsor, legibility for an HTA assessor who cannot run it fits

[1] HTACG JCA dossier guidance (28 Nov 2024), App. B.2 reporting checklist for evidence syntheses: software and version, code for non-standard methods, input data and derivation, MCMC chains and convergence, digitisation method, model and covariate selection, overlap and effective sample size. [2] HTACG guidance on multiplicity, subgroup and sensitivity analyses (10 Jun 2024): for each analysis, prespecified or not, multiplicity control, and by whom it was deemed necessary. ARS v1.0 classes and terms checked against the published LinkML model (19 Apr 2024). “Fits”: covered as is. “Stretch”: a mechanism exists; the vocabulary does not.

What ARS Has No Word For Yet

Criteria for execution

“PT only if incidence ≥ 5%.” “Subgroup curve only if the interaction is significant.” The rule tests a result, not a record, so no where-clause can express it. The template decides what is shown by default.

Record the rule, or only its outcome?

Ingredient provenance

A digitised curve, an event rate copied from a publication, an expert’s prior. ARS can name a source document; it cannot say how the ingredient was made. The first JCA’s error lived exactly here.

Record the recipe for the ingredient too?

Where a number goes

ARS knows a result’s origin and its display. A survival parameter that feeds a cost-effectiveness model has a destination, not a table cell.

Record where a number goes, or only where it came from?

Environment and change

One field for the language and version. Version numbers that say which, never what changed. A delta dossier is a diff nobody can compute.

Record the environment nobody re-runs? Record the diff?

Record as much as possible? How much is too much, and for whom? Each field on this slide costs a statistician time now; the benefit lands later, and often with someone else. That trade is the discussion.

Who Asks for Data and Code Today

Body Asked for today Direction of travel
FDA SDTM, ADaM, define.xml required; source code for ADaM and primary and secondary efficacy outputs expected [1] Dataset-JSON consulted in 2025, not yet adopted; the Data Standards Catalog lists Define-XML only, no ARM, no ARS [2]
EMA Nothing required for a marketing application. Voluntary pilot since Sept 2022, 13 procedures: SDTM, ADaM, programs, ARM “highly recommended” [3] Pilot extended “until further notice”; the network workplan plans a transition to systematic submission of clinical study data in 2026–2028, alongside the revised EU pharmaceutical legislation [4]
EU JCA Documents per PICO; no individual patient data; code only for non-standard methods, and in practice as PDF [5] AI principles for dossiers, July 2026: every AI-assisted step documented, prompts kept, a human responsible [6]
G-BA, NICE G-BA: full CSRs, readable code when deviating from standard software, code plus input data for indirect comparisons. NICE: a fully executable economic model with full code access [7,8] G-BA template (Nov 2025) cross-references the EU dossier; NICE position statement on AI in evidence generation (Aug 2024) [7,9]

Nobody asks for ARS: not one regulator, not one HTA assessor. And no HTA assessor has the ADaM. FDA can re-execute a sponsor’s analysis and EMA is heading there; no HTA body can, by design. So “reproducible” cannot mean at the G-BA what it means at the FDA. What is left is legibility: what was done, for whom, against what, at which cut.

[1] FDA Study Data Technical Conformance Guide v6.2.1 (Jun 2026) §4.1.2.10. [2] FDA Data Standards Catalog v11.0 (Mar 2025); Federal Register, 9 Apr 2025 (Dataset-JSON request for comments). [3] EMA, clinical study data pilot Q&A (EMA/658116/2022, rev. 20 Jan 2026) §3.3–3.4; NDSG annual report 2025 (Jan 2026). [4] EMA Network Data Steering Group workplan 2026–2028 (Feb 2026); EMA Programming Document 2026–28. [5] HTACG JCA dossier guidance (28 Nov 2024) App. B.2, D.3; published JCA dossiers Jun–Sep 2026. [6] HTACG, general principles on the use of AI in JCA dossiers (adopted 15 Jul 2026). [7] G-BA Module 4 template (18 Nov 2025) and dossier format document (8 Apr 2026). [8] NICE technology appraisal manual PMG36 §5.5.15 (updated 31 Mar 2026). [9] NICE ECD11 (15 Aug 2024).

Where This Leaves Us

Re-use is highest in HTA

One recipe, many diners, many cuts. A machine-readable recipe pays back on every re-cook, and nowhere is it re-cooked more.

A shared standard is what makes shared tooling possible

Thousands of records to navigate, trace and diff. Sponsors need that for QC, HTA assessors to find what they asked for. Nobody builds a viewer for an in-house format.

Nobody has asked, and every field has a cost

No regulator or HTA assessor names ARS. Each field costs a statistician time now; the benefit lands later, and often elsewhere: a later dossier, a partner, an HTA assessor.

What would we want to achieve, and what would it cost to capture? Reproducibility, efficiency to reproduce, or legibility for a reviewer and for a machine: pick the goal before picking the metadata. Over to the discussion.

Discussion and Q&A

From Structured Results to Trusted Evidence

ARS can make an analysis computable. What makes it usable, admissible, and trusted beyond the team that produced it?

01 · Efficiency ↔︎ readability

Can we accelerate production without making the result harder for a statistician, reviewer, or model to inspect?

02 · Submission ↔︎ evidence

Should a regulatory or HTA submission carry the ARS/ARD, the rendered output, the execution environment, or all three?

03 · Automation ↔︎ validation

Does machine-readable metadata reduce validation effort, move it earlier, or create a new object that must be validated?

04 · Results ↔︎ decisions

Which PICO, comparator, estimand, data cut, uncertainty, and presentation details must be machine-readable for an HTA assessor to challenge the result?

The question for the room: where should the standard stop, and what governed layer should come next?

Thank You