Overview & Architecture
EMA Forge is an open-source browser-based toolkit for designing mobile ecological momentary assessment (EMA) studies and handing them off to a deployment in the researcher's own Cloudflare account. In one workflow, you can arrange surveys and branching, add optional phone-camera tasks, preview the participant experience, deploy the study with private response storage, and monitor or export responses. The Builder needs no account and does not upload your draft study to EMA Forge.
The key workflow is from idea to a working study link: a participant opens the study in their phone browser without installing an app, while the research team manages its own deployment and data. The guided Cloudflare setup can take minutes for a draft; reviewing measures, consent, costs, device support, and institutional approval takes additional work. Phone physiology tasks are experimental and require study-specific validation.
The workflow has four parts:
You do not host the Builder yourself. Go to emaforge.keeganwhitacre.com, design a study, download the prepared HTML, and install it in the Cloudflare host provisioned in your account. The Builder runs locally in your browser; the deployed Worker and R2 bucket receive responses. Hosting providers may process request metadata, and optional place lookups or SMS involve additional providers.
Building requires no EMA Forge account. The recommended deployment uses your Cloudflare account and R2 billing activation; optional Twilio messaging carries separate charges. Manual static hosting and researcher-approved data-return routes are available as advanced options.
Why Serverless?
The guided deployment provisions a study-specific Worker and private R2 bucket in the researcher's Cloudflare account. You do not maintain an application server, while still documenting these services and the data flow in your protocol. The independent static-host route remains available for teams with an approved receiver or manual data-return process:
- Guided Cloudflare handoff. Install a prepared study in a Worker with its own private response bucket and protected Study Admin. Use a separate deployment for each study.
- Researcher-controlled exports. Inspect live response summaries and download response and dispatch logs from the protected deployment. The independent hosting route can instead use a researcher-approved receiver or manual return.
- Phone-browser participation. Participants open links without an app installation. Test connectivity, supported devices, submission acknowledgments, and recovery before recruitment.
- Inspectable runtime. The prepared study HTML includes the participant-side code and configuration. Researchers can inspect what runs on the phone and document external providers separately.
Tradeoffs to be honest about (see Known Limitations):
- The recommended Cloudflare deployment includes a private study dashboard; a standalone static export does not, and data are only visible after files are returned to the researcher.
- SMS still requires the researcher's own Twilio account and sender. The Cloudflare study host provides scheduling, signed callbacks, opt-out updates, and an auditable dispatch log (see Twilio Integration).
- The static host can see request metadata (IP, timestamp of page load). This is the same constraint as any web page — it's not a data transmission, but it's not nothing.
Feature Status Matrix
Keep this in front of you when planning a study. Anything marked BETA has shipped but hasn't yet been through external pilot validation; anything marked WIP or PLANNED is not yet usable.
| Feature | Status | Notes |
|---|---|---|
| Study Builder (questions, schedule, theme) | STABLE | Schema v2.0.0. |
| Single-file & static-bundle export | STABLE | Both include config.json. |
| Onboarding / consent flow | STABLE | Rich-text consent, progress bar. |
| EMA measures (slider, choice, multi, text, number, instructions, body map) | STABLE | Instruction screens create automatic boundaries; body maps store stable region IDs. |
| Affect Grid (valence × arousal) | STABLE | Stored as {valence, arousal} in [−1, 1]. |
| Skip logic (compound AND/OR) | STABLE | Rules on any prior non-multi question. |
| Phase sequencing (multi-task windows) | STABLE | Arbitrary ordered EMA/Task/HR steps. |
| Conditional tasks (e.g. run ePAT only if HR > 80) | STABLE | See Conditional Tasks. |
| Session locking + crash recovery | STABLE | LocalStorage per-phase resume. |
| Analyze workspace (local import, simulation, CSV export) | STABLE | Runs locally in browser. Scheduled completion is shown only when a schedule manifest exists. |
| Heart-rate capture question type (PPG) | BETA | Requires rear camera + torch. iOS works; desktop no. |
| ePAT task (heartbeat perception) | BETA | Shares PPG core with HR capture. |
| Response latency in Analyze | BETA | Shown when a session includes a notification-to-open duration; imported files without delivery timestamps remain explicitly unavailable. |
| Cloudflare-native Twilio delivery | BETA | Configured through the deployed Study Admin; see Twilio Integration. Automated tests pass; live multi-participant pilot remains required. |
| Webhook auto-upload on session complete | BETA | Use a receiver that returns a readable JSON acknowledgment; confirm data appears in storage. |
| IAT (Implicit Association Task) | EXPERIMENTAL | Development module. Procedure, scoring and mobile timing need independent validation. |
| Stroop / additional task modules | PLANNED | The task system is built to be extended. |
System Requirements
For you (researcher)
- Any modern desktop browser. Chrome, Edge, Firefox, or Safari (from 2022 onward). The Builder and Analyze run entirely in your browser at emaforge.keeganwhitacre.com.
- A way to host your exported study. GitHub Pages is the path of least resistance and free for researchers. Any institutional static-file server works too. (See Hosting Your Study — it is genuinely ten clicks.)
- A way to send participants links at the right times — an SMS service, institutional email scheduler, Twilio, or similar. EMA Forge generates the links; something else delivers them.
- R or Python (optional). Analyze handles local inspection, data-quality review, and long-format export in the browser. Use R or Python for confirmatory and model-based analysis.
For participants
- iOS Safari 14.5+ or Chrome for Android 90+ (roughly anything from 2021 onward).
- For the HR / ePAT modules only: a rear-facing camera with a controllable torch (flashlight). All modern iPhones and most Android devices qualify; tablets without rear cameras do not. Desktop browsers will refuse to launch these modules.
- A working browser link. That's it — no app install, no account.
iOS quirk worth flagging during onboarding: Camera/torch access on iOS requires the page to be launched from Safari (not from inside the Messages app preview or from an SMS link-preview bubble). Your participant instructions should explicitly say "tap the link to open it in Safari" for studies using HR or ePAT.
Quick Start — 5 minutes to a working draft
Go from an idea to a hosted web EMA study with surveys, branching, and optional camera-based physiology. Builder and Analyze need no installation or coding; Cloudflare deploys the participant study, private response storage, and study admin into your account. A draft can take minutes; review consent, measures, device support, and data handling before enrolling anyone.
- Go to emaforge.keeganwhitacre.com and click Builder.
- Click Open protocol library and choose Daily Rhythm & Sleep, or start from the default draft and build your own ordered survey/task flow.
- Replace the consent placeholder with approved study language. Review the questions, sessions, task burden, and device requirements rather than deploying a demonstration protocol unchanged.
- Open Review & Deploy, resolve blocking checks, and select Download prepared study.
- Select Deploy to Cloudflare. Replace the masked example
ADMIN_TOKENwith your own 32+ character password. Twilio is optional and can be connected inside Study Admin later. Copy the resultingworkers.devaddress; EMA Forge accepts it with or withouthttps://. - Paste the address into EMA Forge and select Open admin. A new deployment opens directly on Install study; upload the prepared file.
- Run the live
/check.html. It shows a plain-language pass or failure and stores the synthetic check separately from participant responses. Then complete a real test session on your phone. - Return to
/admin, open Analyze, and select Download analysis file. Open detailed Analyze and import that single.jsonfile to inspect observed sessions and item responses. Completion remains unavailable without scheduled prompt events.
That round trip tells you whether your device can open sessions and whether your data is landing where you expect. Always do this before enrolling a participant.
Please use the hosted version at emaforge.keeganwhitacre.com rather than cloning and self-hosting the Builder. The hosted site is actively maintained and always on the current schema version. Everything happens in your browser — nothing you type into the Builder is transmitted to the server, so "hosted" here just means "the files are served from one canonical place." You'll still self-host the exported study (see Hosting Your Study), because that's the file participants actually open.
Protocol Library
The standalone Protocol Library lets you browse descriptions, burden, device requirements, validation status, and linked evidence before opening an item in the Builder. It provides three kinds of reusable building blocks:
- Complete protocols replace the current draft and demonstrate end-to-end study flows.
- Question packs add a new survey step to the first session while safely remapping question, branching, and response-piping IDs.
- Physiology task presets configure an existing built-in task such as ePAT or heartbeat counting and add it to the first session.
Every item is versioned, inspectable JSON. A local community item can be installed with Install library JSON. The Library's guided creator can copy the current browser-saved study, validate and download the contribution JSON, and open a prefilled Submit for review request. Public items remain human-reviewed and are never published automatically. Established measures should disclose the exact version, permissions, timeframe, scoring, validation status, and any adaptations. Library JSON cannot execute uploaded JavaScript or introduce an unreviewed task engine.
A library item is a starting point, not an approved protocol. Verify licensing, validation evidence, population fit, burden, consent, and safety procedures before use.
Study Tab
Global study configuration. Set here:
- Study name & institution — appears in participant app title and CSV metadata.
- Theme & accent color — OLED (pure black, recommended for battery), dark, or light. Accent color inherits into the participant app.
- Response format — CSV (one row per answered question) or JSON (full structured payload). Both are downloaded; JSON is always included when a session contains high-density signal data (PPG/ePAT).
- Webhook URL (optional) — if set, session data is auto-uploaded to this URL instead of being manually downloaded by the participant. See Auto-upload with Webhooks for setup.
- Greetings — per-window header text ("Good Morning", "Check-In", etc.).
- Completion lock & crash recovery — see Runtime Behavior.
Onboarding & Consent
The Onboarding flow runs exactly once per participant, on their first link
(?id=N&session=onboarding, Day 0). It covers:
- A progress-barred multi-screen intro.
- A rich-text consent screen that requires scroll-to-bottom before the "I agree" checkbox activates. Consent text accepts HTML so you can paste in your IRB-approved language with headings.
- An optional schedule-sanity screen letting the participant confirm they'll be available during each window.
- For sessions that include ePAT: a two-phase practice (tone-to-tone, then tone-to-heartbeat). The ePAT measurement itself runs at its scheduled session, never during setup.
If onboarding is disabled, Day 0 is skipped entirely and participants open their first scheduled session.
Measures
Measures is the single entry point for survey items and physiology. Choose a session, then add a rating, response item, PPG heart-rate capture, ePAT, or heartbeat-counting task. The participant-flow summary shows survey blocks and physiological tasks together; Schedule controls their exact order. Survey items appear as expandable cards and update the participant preview as you type.
Question types
| Type | Stored as | Notes |
|---|---|---|
slider |
number | Min, max, step, unit, two anchor labels. |
choice |
string | Single-select from an option list. |
checkbox |
array of strings | Multi-select. Serialized as a;b;c in CSV. |
text |
string | Free-form text input. |
numeric |
number | Number input with keypad on mobile. |
affect_grid |
{valence, arousal} |
2D tap target. Both axes in [−1, 1]. Serialized as valence;arousal in CSV. |
heart_rate |
{bpm, sqi, ibi_series} |
Camera PPG for a configurable duration. The BPM number is what participates in skip logic and conditional tasks. The "Show BPM to participant" toggle controls whether the number is rendered on screen - turn it off before HCT or any other task taht could be biased by knowing the captured HR. See HR Capture. |
place_context |
{status, indicators, dataset, version} |
Optional current-location lookup using five selectable EPA/Census measures in one Cloudflare lookup; researcher-provided study-area polygons can run on-device. No coordinates or area IDs are saved in responses; CSV stores the result as JSON. |
page_break |
— | Splits questions onto separate screens. Not a question per se. |
Place context lookup (optional)
Choose Automatic EPA walkability for the recommended Cloudflare deployment. Researchers do not need to provide a dataset. After a participant taps Use my location and grants browser permission, the phone sends coordinates in a POST request to the study Worker, which queries EPA’s National Walkability Index mapping service. The Worker returns one category and does not write coordinates or EPA’s block-group ID to R2. EPA and Cloudflare receive the lookup request and may retain service metadata under their own policies. This is a historical, area-level built-environment measure, not a live walkability or pollution reading. If EPA is unreachable or the device position is too imprecise, the response contains a missing status. The independent static-host export does not provide this online lookup endpoint.
Choose Automatic Census urban/rural to query the U.S. Census Bureau’s 2020 Census block urban/rural flag through the same study Worker. You can also select three EPA Smart Location Database v3 measures: population density (2018 estimate, people per developable acre), distance to the nearest transit stop (from the block-group population-weighted centroid, not the participant), and share of households without a car (2018 estimate). These are historical context measures, not observations of the participant, transit service today, or an individual's means. Select any combination of five automatic measures in one question: the participant grants permission once, and the Worker sends coordinates to EPA and/or Census as needed. The three Smart Location measures share one EPA request. A response records selected categories and source versions; partial results include an indicator_statuses entry for each, and missing values remain missing. No dataset upload is needed. No response saves coordinates or a block ID. There is no automatic suburban classification. Unavailable services, positions outside coverage, and imprecise locations have distinct missing statuses. This lookup sends coordinates to Cloudflare and each selected provider, which may process request metadata. Use the on-device option if that transfer is inappropriate for your study.
EMA Forge uses fixed descriptive bands for the new EPA measures; they are not validated clinical cutoffs. Population density is low under 1, moderate from 1 to under 10, and high at least 10 people per developable acre. Transit distance is near at 400 meters or less, intermediate above 400 through 1600 meters, and far beyond 1600 meters. Car-free household share is low under 5%, moderate from 5% to under 20%, and high at least 20%. Review coverage and source definitions before using bands as research variables. Even without stored coordinates, combining multiple categories, response times, and repeated observations may disclose a participant's home or routine; select only what your protocol needs. Area-level car ownership must never be treated as the participant's ownership or income.
Choose On-device study-area data if coordinates must stay out of lookup requests. In this mode, the researcher prepares a licensed dataset, and matching happens in the participant’s browser. None of the modes stores coordinates in EMA Forge response records. Only the on-device mode can include a researcher-defined pollution or deprivation band; its source, period, and thresholds need to be documented. An area-level vulnerability indicator describes a place, not a participant's income or identity.
Adding this measure opens a setup notice about consent, location privacy, source licensing, and accuracy. For the on-device mode only, upload a small GeoJSON FeatureCollection (up to 2 MB) containing study-area Polygon or MultiPolygon features. Include metadata fields name, version, source_url (HTTPS), and method describing how source values became categories. Put only permitted categorical values inside each feature's properties.indicators: walkability, pollution, deprivation (very_low / low / moderate / high / very_high), or urbanicity (urban / suburban / rural). Example feature: {"type":"Feature","properties":{"indicators":{"walkability":"high"}},"geometry":{"type":"Polygon","coordinates":[[[...]]]}}. This is a format example, not an index or a ready-to-use data source. The study-area polygons are embedded in the publicly delivered study file, so upload only data you have the right to redistribute.
A pollution category needs a specific pollutant, geographic source, reference period, and threshold method; EMA Forge does not invent or automatically calculate those values. Neighborhood Atlas ADI has separate data-use terms; obtain permission before packaging it. A browser permission prompt appears after participants tap Use my location. Declines, denied access, locations outside coverage, and uncertain matches have distinct statuses. Preview never asks for actual location. Location services on the device may use their own providers, and derived categories can still reveal sensitive context when paired with participant ID and response time. Review study consent and institutional requirements before collecting these data.
Per-question settings
- Show in these survey steps: choose the exact session steps where this question appears. Select more than one step when you intentionally want to repeat a question.
- Required: blocks advancement until answered.
- Skip Logic: compound AND/OR conditions on any prior question. Operators:
eq, neq, gt, gte, lt, lte, includes. When a condition evaluates false, the question is silently skipped and not recorded. - Earlier-answer personalization: in a later question, use Insert earlier
answer, choose the source item, and set fallback wording for participants who did not answer it.
EMA Forge inserts the answer at the cursor, formats multi-select responses as readable prose, and checks
that the source really appears earlier in every session where the follow-up is used. In JSON, the stored
form is
{{q_id|fallback wording}}; the builder handles this syntax for you.
Skip logic semantics — edge cases worth knowing
Conditions reference the raw response value, which means:
- For sliders, comparisons are numeric.
q1 gt 70works as expected. - For single-choice, comparisons are string equality against the option label (not index).
- For checkbox (multi-select), comparisons use
includes.eqon a multi-select will nearly always be false because the stored value is an array. - For Affect Grid, skip-logic is disabled in the UI (the value is a compound object, not a scalar).
- For Heart Rate, comparisons hit the
bpmfield directly, soq_hr_1 gt 80does what you'd expect. - If a question referenced by a condition was itself skipped or not-yet-presented, the rule evaluates false. Design accordingly.
Sessions & Scheduling
New studies start with one daily ePAT session and no survey questions. Choose the study length and active days, then add sessions as needed. Each session has a time window and an ordered participant flow. Add survey questions, PPG heart-rate capture, ePAT, or heartbeat counting from Measures or directly from a session. Adding a task enables its module; advanced controls stay under Task settings. Survey questions are optional.
Days & windows
- Study length: 1–365 days.
- Days of week: restrict to weekdays only, etc.
- Windows: each has an ID, a display label, and a start/end time. The ID (e.g.
morning) is what becomes thesession=URL parameter.
Participant flow
Drag the cards or use their arrow buttons to change the order. Add as many separate survey steps as your
protocol needs. Measures assigns each survey item to specific steps. Each card has an Options control
for its name or task selection and conditional task rules. The export stores this order as
phase_sequence, with question_ids on every survey step.
The available steps are:
| Step kind | Meaning |
|---|---|
ema |
An independent set of questions selected for this step. You can add several survey steps in any order. |
task |
An enabled task module: ePAT, heartbeat counting, or experimental IAT. Optionally gated by a condition. |
heart_rate question |
Camera-based heart-rate capture in its own survey step or alongside other questions. |
Steps run top to bottom. For example, add PPG heart-rate capture, then ePAT, then a new survey step. Questions only repeat when you select multiple survey steps for the same question.
Response window timing: the expiry_minutes setting is enforced at runtime
by comparing against a t= URL parameter (a millisecond timestamp added by your
SMS/scheduler). If t is absent, links never expire — convenient for piloting, but turn this
on before real enrollment.
Task settings
Add a task to a session from Measures or Schedule to enable it automatically. In Task settings you can adjust its settings or disable a module; a disabled task still assigned to a session will block export until re-enabled or removed.
Currently available: ePAT, HCT, and experimental IAT. No camera is required for the IAT — it is the only task module that works on any smartphone without hardware constraints.
Live Preview
The Builder contains a live participant app and the top-bar Preview button opens it in a
full-screen iPhone-sized frame. Preview mode (__PREVIEW_MODE__ = true) disables link expiry and
session locking so you can repeat a session. It also clearly simulates camera access, flashlight checks,
PPG, calibration, ePAT, and HCT values; the Builder preview never asks for sensor permission.
Click Reset to restart the rendered session without changing the study draft. Preview is for flow and usability checks, not device compatibility or measurement validation; those checks still require a deployed HTTPS study on a supported phone.
The compact Simulated preview ⓘ badge expands on demand, so the notice remains available without displacing the participant screen.
Export Options
| Option | What you get | When to use |
|---|---|---|
| Prepared Cloudflare study RECOMMENDED | One installable HTML file with the same-host /submit receiver already embedded. |
Normal studies. Deploy the EMA Forge Cloudflare template, then upload this file in Study Admin. Do not paste the Worker hostname into the manual webhook field. |
| Single HTML file | One self-contained .html with config, JS, and CSS inlined. |
Pilots, QA, or emailing a direct file to one participant. Config is baked in — any edit means re-exporting. |
| Static-hosting bundle | A .zip containing index.html, config.json,
check.html, css/, and js/. |
Real deployment. Because config.json is separate, you can fix a typo in a
question without re-exporting. Just edit the JSON on the host. |
Preflight validation runs on every export. Blocking errors include placeholder consent, duplicate IDs, empty or unreachable EMA steps, broken conditions, unavailable task modules, invalid timing, malformed webhooks, and unsafe module settings. Research warnings — such as manual-only data return, disabled in-app consent, experimental IAT use, or unpiloted physiology — require explicit confirmation. Deployment files also reject placeholder, local, non-HTTPS, or query-bearing host URLs.
Hosting Your Study
This section is about hosting the exported study bundle — the thing participants will open. You are not hosting the Builder itself (that already lives at emaforge.keeganwhitacre.com).
The simplest complete option is the Cloudflare study host in Review & Deploy: it serves the participant app, stores responses privately, and provides the protected Admin workspace in one deployment. The GitHub Pages path below is an advanced static-host alternative and still needs a separate approved receiver for automatic response return.
- Create a free GitHub account if you don't have one.
- Create a new repository (e.g.
my-ema-study). Make it public or private — both work with GitHub Pages on free accounts. - Click Add file → Upload files, and drop in the entire contents of your extracted Export
zip (the
index.html,config.json, and thecss/jsfolders). - Go to Settings → Pages. Under "Source," choose main branch, root folder.
- Wait a minute. Your study is now live at
https://<your-username>.github.io/my-ema-study/.
Enter that URL in Review & Deploy. If automatic return is enabled, open the hosted
check.html, send its synthetic test, and confirm the same ID exists in receiver storage. Then
complete one full participant session and inspect that record before generating enrollment links.
Prefer the command line, or deploying via another host?
Any static host works unchanged — Netlify, Vercel, Cloudflare Pages, S3 + CloudFront, or an Apache/Nginx directory on institutional infrastructure. No build step, no server runtime.
The git-native GitHub Pages recipe:
# From your extracted bundle folder
git init
git add .
git commit -m "study v1"
git branch -M main
git remote add origin https://github.com/your-lab/your-study.git
git push -u origin main
# Then in GitHub: Settings → Pages → Source = main branch, root
# Your study is live at https://your-lab.github.io/your-study/
Participant Routing
Because there is no backend, the participant's session is entirely determined by the URL they open. Four query parameters drive this:
| Parameter | Required | Description |
|---|---|---|
id |
yes | Participant identifier (any string, usually numeric). |
day |
yes (except onboarding) | Study day, typically 1–N. |
session |
yes | Window ID defined in the Schedule tab, or the literal string onboarding for Day 0. |
t |
recommended | Millisecond Unix timestamp of when the link was sent. Used for expiry enforcement. Your scheduler should inject this at send-time. |
force |
no | When set to 1, overrides session locking. For researcher use only (QA, rescues). Do not
put this in participant links. |
https://your-lab.github.io/study/?id=104&day=2&session=morning&t=1732812000000
Generating Links in Bulk
The Builder's Review & Deploy tab takes a base URL and a participant-ID range, and emits a CSV
with one row per (participant × day × session). Each row includes a Phase_Sequence column so
you can eyeball at a glance what each link will do.
Participant_ID,Day,Active_Weekdays,Session,Phase_Sequence,Send_Window_Local,URL_Template
104,0,Any,Setup,Onboarding,Researcher scheduled,https://.../?id=104&session=onboarding
104,1,Mon|Wed|Fri,Morning,Survey questions → ePAT → Reflection,08:00-10:00,https://.../?id=104&day=1&session=w1&t={sent_at_ms}
...
This CSV is the glue between EMA Forge and your delivery mechanism of choice. It now emits a
URL_Template containing {sent_at_ms}; your delivery system must replace that token
with the millisecond Unix timestamp at send time. Active_Weekdays and
Send_Window_Local make the scheduling rules explicit rather than silently implying every row
should be sent.
Sending Prompts
You have two paths. Both use the same participant, day, window, and send-time routing model.
- Bring your own mechanism. Drop the routing CSV into an institutional email merge, a
Qualtrics contacts list, a Power Automate flow, or any SMS tool that can schedule messages from a CSV. You
handle timing and
t=injection. - Use Cloudflare Study Admin. BETA Add the roster
and connect Twilio inside the deployed study. The Worker handles timezone math, dedupes sends, records
delivery callbacks, and enforces
t=expiry. See Twilio Integration.
Session Lifecycle
When a participant opens a link, the runtime:
- Parses URL parameters.
- Checks expiry (if
tandexpiry_minutesare both set). - Checks the completion lock — has this
(id, day, session)tuple already been submitted on this device? - Checks for an in-progress resume state for this session.
- Renders the first step of the session (usually a pre-task EMA block).
- On each phase completion, writes the phase's responses to local storage and advances.
- When the last phase finishes, the full session payload is assembled and the "Save Local Copy" screen appears. Tapping it triggers a file download.
Expiry & Grace Period
- Expiry window (default 60 min): after this many minutes past
t, the link renders a "Link Expired" screen and the Start button is disabled. Recorded as a missed ping by the Dashboard. - Grace period: a participant must begin before expiry. Once started, the exported runtime
allows completion until
expiry + grace. Crossing that hard deadline records and delivers the partial session withstatus: expired_in_progressandendedReason: response_window_elapsed.
Session Locking
By default, a participant who completes a session and then re-clicks the same link sees a "Session Complete — thanks, come back at the next prompt" screen. This prevents accidental double-submissions.
Researchers can override with ?force=1 appended to the URL (useful for QA walkthroughs or for
rescuing a participant whose submission didn't land). Do not include force in
participant-facing links.
Crash Recovery
If a participant's browser crashes mid-session (tab killed by iOS memory pressure, phone restart, app switch past timeout), reopening the same link will:
- Restore all completed steps (surveys, finished task trials, etc.) — that data survives.
- Restart the current phase from the beginning. Partial data within the phase is discarded to avoid ambiguity.
This is intentionally conservative. A half-finished phase with an unknown number of missing questions is worse than a clean re-run.
Schema Versioning & Migrations
Every study carries a schema version (currently 2.0.0). When the Builder
loads a saved study — whether from your browser's local storage or a backup file you imported — it checks
that version:
- Equal → load as-is.
- Older → refuse to load. Version 2 assigns questions directly to steps; start a new study or retain an older backup outside this Builder.
- Newer → refuse to load. A newer schema might include fields this runtime doesn't know how to preserve; loading would silently corrupt them. Export a backup from the newer Builder or reset.
This is worth knowing when rolling out an EMA Forge update mid-study: finish in-progress waves on the old version, then upgrade.
What Happens at Session End
When the last phase completes, one of two things happens depending on whether you've configured a webhook (Study tab → Webhook URL):
Manual return (when no receiver is configured)
The participant sees a "Save Local Copy" button. Tapping it downloads one or two files to their device:
ema_data_[ID]_[DAY]_[SESSION].csv— long-format, one row per question answered. Default output when the session contains only standard EMA responses.ema_data_[ID]_[DAY]_[SESSION].json— full structured payload, included automatically whenever the session contains high-density signal data (PPG samples, ePAT trial-level data). Written alongside the CSV, not instead of it.
The participant returns these files via whatever channel your IRB approved (secure email, REDCap file upload, institutional SFTP). Analyze accepts Cloudflare NDJSON, a folder of JSON files, or compatible long-format CSV; a CSV-only workflow is also viable for simpler studies.
Automatic return
The prepared Cloudflare study uses same-host /submit automatically. For an independent host,
set an approved HTTPS Webhook URL. The participant sees an upload confirmation only
after a readable success receipt. Otherwise "Save Local Copy" appears. Browser connectivity, receiver CORS,
and actual server storage must be checked with a dummy submission before collecting participant data. See
Auto-upload with Webhooks.
CSV Schema (Analyze export)
When you use Export CSV in Analyze, you get one row per
(session, question) pair in long format. The columns:
| Column | Description |
|---|---|
data_source |
observed for imported records or synthetic for simulator output. |
participant_id |
From the URL ?id=. |
day |
From the URL ?day=. |
session_id |
Unique per-session identifier generated by the runtime. |
window_id |
Matches the Schedule-tab window ID (e.g. morning). |
window_label |
Human-readable window label. |
block |
Survey step ID or blank. Identifies the exact survey step that presented this question. |
session_started_at |
ISO-8601 timestamp of session start. |
session_submitted_at |
ISO-8601 timestamp of final session save. |
phase_started_at |
ISO-8601 timestamp — when the block containing this question began. |
phase_submitted_at |
ISO-8601 timestamp — when the block was submitted. |
question_id |
Stable ID (e.g. q1, q_hr_1). |
question_text |
The question text at the time of export. |
question_type |
One of the types listed in Measures. |
presentation_order |
1-indexed order the question was shown in, accounting for skip logic. |
response_status |
answered, unanswered, or skipped_condition. |
skip_reason |
Machine-readable reason when branching prevented presentation. |
response_value |
The raw response, serialized (checkbox as a;b;c, Affect Grid as
valence;arousal).
|
response_numeric |
Numeric cast of the response if possible; blank otherwise. Convenient for sliders. |
response_latency_ms |
Milliseconds from phase start to this response. |
JSON Schema (per-session)
The raw JSON payload is what Analyze reads and is the source of truth. Shape (abbreviated):
{
"participantId": "104",
"sessionId": "s_a1b2c3",
"day": 2,
"type": "ema_with_task",
"status": "complete",
"startedAt": "2026-04-22T08:03:11Z",
"completedAt": "2026-04-22T08:06:48Z",
"data": [
{
"type": "ema_response",
"block": "s_baseline",
"windowId": "w1",
"startedAt": "...",
"submittedAt": "...",
"presentationOrder": [["q1", "q2"], ["q3"]],
"responses": {
"q1": { "value": 72, "respondedAt": "..." },
"q2": { "value": "Working", "respondedAt": "..." },
"q3": { "value": {"valence": 0.4,
"arousal": -0.2}, "respondedAt": "..." },
"q_hr_1":{ "value": { "bpm": 74,
"sqi": 0.82,
"ibi_series": [810, 795, ...] },
"respondedAt": "..." }
}
},
{
"type": "epat_response",
"trials": [ { "trial": 1, "phase_ms": 342, "confidence": 3, "sqi": 0.91 }, ... ],
"summary": { "valid_trials": 18, "mean_abs_phase_ms": 218, "..." }
}
]
}
Analysis in R
Because the Dashboard's CSV export is long-format, getting from a folder of per-session files to a tidy dataset is short. A typical starter pipeline:
library(tidyverse)
library(jsonlite)
# 1. Ingest a folder of per-session JSON files
files <- list.files("data/", pattern = "\\.json$", full.names = TRUE)
sessions <- map(files, ~ fromJSON(.x, simplifyVector = FALSE))
# 2. Flatten to long format (or just use the Dashboard's master CSV)
df <- read_csv("ema_master_dataset_2026-04-22.csv")
# 3. Compliance per participant
compliance <- df |>
distinct(participant_id, day, session_id) |>
count(participant_id) |>
rename(sessions_completed = n)
# 4. Mean mood by time-of-day, handling the long format
mood <- df |>
filter(question_id == "q1") |>
group_by(participant_id, window_label) |>
summarize(mean_mood = mean(response_numeric, na.rm = TRUE),
sd_mood = sd(response_numeric, na.rm = TRUE),
n = n(),
.groups = "drop")
# 5. Affect Grid (stored as "valence;arousal" in response_value)
affect <- df |>
filter(question_type == "affect_grid") |>
separate(response_value, into = c("valence", "arousal"),
sep = ";", convert = TRUE)
Analyze
dashboard.html is a local analysis and protocol-simulation workspace. Imports and synthetic
generation run in your browser; EMA Forge does not upload the files to a central analysis service.
What to feed it
Analyze accepts five input shapes. You can mix compatible files in a single import — the parser sorts out which files are which.
- One analysis file from Study Admin (recommended). In the deployed study's
/adminpage, open Analyze and select Download analysis file. It includes the installed study's question configuration and all exported responses in one JSON file. Import just that file into detailed Analyze. Keep it in an institutionally approved location: it contains participant responses. - A Cloudflare response export. From the deployed study's
/adminpage, select Download responses and import the resulting.ndjsonfile directly. Each line is one complete, lossless session envelope. Import multiple export parts together when the admin page indicates that more than one part is available. - A folder of JSON files. One per-session
.jsonas produced by the participant app, plus a singleconfig.json(the same one inside your exported study). This is the right shape when participants return files manually, or when you've downloaded them out of REDCap/institutional storage. - A CSV from the Google Sheets webhook. BETA If
you're using the Google Apps Script webhook recipe, your sheet has a
Raw JSONcolumn holding each session's payload. Export the sheet as.csv(File → Download → CSV) and import it alongside yourconfig.json. The parser reads each row'sRaw JSONcell as a full session payload. - A long-format EMA Forge CSV. Files with
session_idandquestion_idcolumns preserve answered, presented-but-unanswered, and condition-skipped states when re-imported.
Study matching. If importing raw responses, select the matching
config.json in the same import. Analyze no longer silently reuses a configuration
cached from a different study. The single analysis file already contains its matching configuration.
Once imported, Analyze renders:
KPIs
- Protocol completion — shown only when an explicit schedule manifest supplies the expected prompts. It is unavailable for response-only imports because missing files cannot prove a missed prompt.
- Delivery records — shown only when delivery events exist. EMA Forge does not infer delivery from a completed response.
- Rapid-session flags — sessions with total duration under 30 seconds. This is a review flag, not an automatic exclusion or a signal-quality verdict.
- Mean completion time — mean duration across observed completed sessions.
Filters & views
- Aggregate vs. Per-Participant segmented view.
- Date range (by study day).
- Exclude rapid sessions (<30s) toggle for sensitivity review.
- Records requiring review — for simulations with a schedule manifest, identifies participants below 60% protocol completion. It stays unavailable for response-only imports.
- Responses — type-aware summaries for numeric, categorical, text, affect-grid, and heart-rate items, plus an exploratory composite-score builder.
- ePAT — trial and session summaries, circular phase visualization, confidence, and signal-quality views.
- Export CSV — flattens all imported sessions into one long-format file (schema above).
Protocol-aware simulation
Select Simulate study to generate a deterministic synthetic dataset from a curated protocol. Choose participant count, study days, prompt completion, item missingness, and a random seed. The generator follows the protocol's ordered survey/task steps and branching, and produces compatible ePAT and heartbeat-counting envelopes when those tasks run.
Synthetic provenance is persistent. Simulated studies and sessions are marked
synthetic in memory; CSV exports add data_source=synthetic and use an
ema_synthetic_dataset_... filename. Simulation demonstrates the end-to-end workflow. It
does not validate a measure, device, task, intervention, or statistical result.
Do not infer missing events. Notification-to-open latency is shown only when imported sessions contain a usable latency value. Response-only imports show it as unavailable. Simulations include synthetic notification latency so the chart can be tested, but those values are not empirical.
Heart-Rate Capture (Camera PPG)
EMA Forge ships a lightweight photoplethysmography (PPG) pipeline that recovers heart rate from the participant's rear-facing smartphone camera, with the torch on as an illumination source. It powers two user-facing features:
- A
heart_ratequestion type you can drop into an EMA block anywhere. - A standalone HR step in a window's phase sequence (identical capture, structural placement differs).
If camera access cannot start, heart-rate, ePAT, and HCT screens offer Try camera again and Continue without measurement. Continuing records structured unavailable-task metadata, preserves completed survey answers, and never invents a physiological value. Researchers should define device eligibility and missing-data handling in the approved protocol before enrollment.
How the PPG pipeline works (click to expand)
The runtime's ePATCore module handles PPG end-to-end. The current pipeline is the v5
architecture; an engineering-level deep-dive of what changed since v4 is in the next collapsible. In
brief:
- Camera acquisition.
getUserMediarequests the rear camera at 30 fps (with 60 fps and "ideal 30" fallbacks) and the torch is engaged viaapplyConstraints({ torch: true }). A hidden<video>element receives the stream; cameras labeled "front", "facetime", "dual", "triple", "ultra", or "tele" are deprioritised so the chosen camera is the standard rear sensor. - Sampling. A hidden
<canvas>draws each video frame and averages pixel intensity over a central region of interest. Both the red and green channels are extracted and tracked separately — the red channel dominates with the torch on, but the green channel is kept available as a backup. Sampling is locked to the camera's actual paint rate (typically 30 Hz); a variable-frame-rate deduplication step prevents the 60/120 Hz display refresh from being mistaken for new data. - Filtering. Each channel is bandpass-filtered with a pair of cascaded 2nd-order Butterworth biquad filters: a high-pass at 0.67 Hz to remove respiratory drift and finger micro-shifts, and a low-pass at 3.33 Hz to remove high-frequency sensor noise. The pass band (0.67–3.33 Hz, i.e. 40–200 BPM) covers the cardiac range. The biquads use Butterworth Q (1/√2) and standard digital biquad coefficients computed at the actual sample rate.
- Signal-quality index (SQI). SQI is computed every 1 s on a 2 s rolling
window of the filtered signal as the peak-to-peak amplitude in the cardiac band. This
directly measures the AC pulse amplitude that beat detection actually requires, and is
scale-comparable across the red and green channels. The runtime exposes per-channel SQI plus the
active channel via
onSqiUpdateCb(sqiRed, sqiGreen, activeChannel). The "Sensor Warning" overlay enforces the floor — if SQI on the active channel drops below the warning threshold, the warning reappears and the participant is prompted to reposition their finger. - Channel selection and failover. Both channels run a full filter + beat-detector pipeline continuously. After a 4.5 s settling window following finger detection (which lets the high-pass transient ring down — its time constant is ~0.24 s, so 5τ ≈ 1.2 s with comfortable margin), per-channel baseline amplitudes are locked. The active channel is initially red. If the active channel's amplitude drops below 40% of its baseline AND the backup channel's amplitude exceeds 25% of its own baseline, sustained for two consecutive 1 s checks (anti-thrash), the runtime switches to the backup. Because both detectors are always running, the new active channel has already learned its thresholds — there is no cold-start blackout on switch. This makes the system robust to drift in finger pressure, ambient lighting changes, and skin tone differences that affect torch saturation differently across channels.
- Beat detection. Each channel feeds an independent instance of a JavaScript port of the WABP (Waveform Analysis for Blood Pressure) algorithm — Zong, Heldt, Moody & Mark (2003) — adapted from arterial pressure to camera PPG. WABP is well-suited because both signals share an upstroke-dominant morphology. After a candidate onset is detected, a quadratic interpolation step refines the peak position in the slope-energy curve to sub-frame resolution. The camera is still 30 Hz, but onset times are no longer quantised to ~33 ms steps — IBIs are computed at finer resolution, which matters substantially for HRV-style analyses where ~30 ms quantisation noise is large relative to physiological variability.
- Dicrotic-notch rejection. PPG morphology contains a secondary peak (the dicrotic
notch) per cardiac cycle that can be falsely detected as a beat, inflating HR by up to a factor of two
when it happens systematically. The runtime maintains a running median of recent inter-beat intervals;
candidate beats whose interval is less than 60% of the expected period are rejected as dicrotic.
Rejection counts are exposed in
diagnostics.dicroticRejectsfor quality auditing. - Output. Each accepted beat is emitted via
onBeatCb({ instantBPM, averageBPM, instantPeriod, averagePeriod, time }).averageBPMis the median of the last 10 inter-beat periods; this median is used in preference to the running mean because it is robust to the occasional missed or extra beat that any peak detector will produce in real-world capture conditions. For HR-question outputs, BPM is reported as60 000 / median(IBIs)over the capture window. The full IBI series is stored alongside the BPM value for downstream HRV analysis, plus the mean SQI as a quality annotation.
This is not a medical-grade measurement. For research where the construct is "the participant's approximate HR at this moment under ecological conditions", the technique is well-established in the digital biomarker literature, and the pipeline above is materially closer to PhysioNet-grade processing than the typical smartphone-PPG demo.
v5 architecture: dual-channel + signal-quality (engineering deep-dive)
This section documents what changed in the signal pipeline between v4 and the current v5 architecture, for researchers replicating the methods or auditing the code. None of this is necessary to use EMA Forge — it's here so reviewers and replicators have an accurate target.
Why dual-channel
A single-channel PPG pipeline has a fundamental fragility: the channel's signal-to-noise depends on the interaction of skin tone, finger pressure, ambient light, and torch brightness in ways that drift across a session. v4 used the red channel exclusively; in practice this worked well in most cases but produced occasional sessions where the red channel saturated (torch + light skin + heavy pressure) or under-illuminated (torch + dark skin + light contact), and the entire session became unscoreable. v5 keeps red as the default but tracks the green channel in parallel and switches if conditions warrant. Importantly, the green pipeline is never "off" — both channels run the full filter and beat detector continuously, so the backup channel has already adapted its thresholds at the moment a switch fires.
Why the SQI metric changed
v4 computed SQI as a perfusion-index (max − min) / mean on the raw camera
channel. This had two pathologies:
- During finger placement, the raw signal ramps from ambient (~0.10) to finger-on (~0.40) inside the
SQI window.
(max − min) / meanover that ramp is dominated by the DC step, not the AC pulse — green could routinely score 14+ versus red's 0.6, making channel selection meaningless. - The threshold was calibrated against torch-saturated red (where raw PI is ~0.01). Green's baseline raw PI is ~0.35. Any threshold tuned for one channel's operating point is wrong for the other.
v5 SQI is the peak-to-peak amplitude of the filtered signal over a 2 s rolling window.
Because the bandpass strips the DC component, this metric measures the actual AC pulse amplitude — which
is what WABP cares about — and is scale-comparable across channels because both feed the same filter.
The raw perfusion index is still computed and exposed in diagnostics for callers that want
it, but no switching decision uses it.
Calibration state machine
After a finger-on transition, the calibration state machine moves from 'settling' →
'locked'. During settling (4.5 s), no channel switching can occur and no beats are
emitted — this gives the high-pass filter time to ring down past its transient. At lock time, each
channel's current filtered peak-to-peak becomes its baseline. From that point onward, failover
thresholds are expressed as fractions of each channel's own baseline (40% drop on active, 25% viability
on backup), which makes them dimensionless and channel-agnostic. After a switch, the new active
channel's baseline is updated to its current value, so subsequent failovers in either direction work
symmetrically.
Sub-frame interpolation
WABP returns the index of the slope-energy local maximum that triggered the detection. With a
30 Hz camera, naive use of this index quantises beat times to ~33.3 ms. v5 fits a quadratic to
the three samples around the local maximum and computes the fractional offset of the true peak:
Δ = ½ · (s₀ − s₂) / (s₀ − 2s₁ + s₂). The reported framesAgo is then
integer_framesAgo − Δ, so beat times resolve at finer than the frame interval. The camera
sample rate has not changed; the beat-time estimate just no longer has 30 ms quantisation noise on
top of the underlying physiological variability. For downstream HRV use this is meaningful — RMSSD on
quantised IBIs is biased upward.
Dicrotic-notch rejection
The dicrotic notch is the mid-diastolic secondary peak in the PPG waveform. WABP's onset criterion can
fire on it under certain morphologies, generating a phantom beat with an IBI roughly half the true
value. v5 maintains a 10-element running buffer of recent accepted IBIs; once it has at least 3 samples,
candidate beats whose interval to the previous accepted beat is less than 60% of the median expected
period are rejected as dicrotic and counted in dicroticRejectCount. The 60% threshold is
conservative — physiological IBI variability (HRV) under normal conditions is well below that ratio.
API contract notes
BeatDetector.setCallbacks({...})swaps callbacks without resetting filter or detector state. This is what allows ePAT and HCT to keep the camera session alive across baseline → trial transitions without a cold-start.BeatDetector.stop()tears down the camera, releases the torch, and resets all internal state. After this, the next call must bestart().- All beat-time stamps use
performance.now(), which is monotonic and unaffected by wall-clock changes. onSqiUpdateCb(sqiRed, sqiGreen, activeChannel)reports both channels every 1 s; the active channel string is one of'red'or'green'.
Test the pipeline on your device: Open ppgtester.html to run a live camera capture and inspect beat
detection, raw waveform, SQI, and IBI output in real time — without needing a full study session. Useful
for validating that PPG works on a target device before deploying, or for checking signal quality in a
given environment.
ePAT — Ecological Phase Assessment Task
The ePAT is a cardiac interoceptive-accuracy task adapted for in-the-wild, mobile-first use. It is directly descended from the Phase Adjustment Task (PAT) and its refined successor PAT 2.0 — see Acknowledgments & Prior Work for primary sources. Participants align an auditory tone with their own felt heartbeats by rotating a dial until the tone "feels like it's landing on" each beat. Their phase offset (in ms, relative to ground-truth peaks detected by the camera PPG) is the dependent measure.
Task flow, trial by trial (click to expand)
- Baseline calibration. 10–20 s of still finger-on-camera capture establishes the participant's current HR and verifies that SQI is above threshold before any trials begin.
- Two-phase practice (optional, default on). First, tone-to-tone alignment (no heartbeat involved — teaches the dial mechanic). Second, tone-to-heartbeat alignment at a slow scaffolded pace.
- Trial block.
trialsvalid trials (default 20), eachtrial_duration_secseconds long (default 30). During a trial:- Live camera PPG streams in the background.
- A tone plays at a predicted time, offset by a randomized phase.
- The participant rotates the rotary dial to shift the tone earlier or later until it subjectively aligns with their beat.
- On "Confirm Timing", the offset is recorded relative to the ground-truth peak from the PPG.
- Per-trial confidence (optional): 1–5 rating of how sure they were the tone matched their heartbeat.
- Body map (optional): after each trial, where did they feel the beat? Chest / fingers / neck / ears / abdomen / legs / head / nowhere.
- Retry budget.
retry_budget(default 30) is the maximum number of attempts. A trial can fail for low SQI, excessive movement, or participant cancelation. Once valid trials hits the target, the task ends; if the budget exhausts first, the task ends with whatever valid trials were collected.
High-precision timing. Audio stimulus scheduling uses the Web Audio API's
AudioContext (not setTimeout), which gives sample-accurate timing across
browsers. This is the difference between a task that has <2 ms jitter and one that has ~20 ms jitter,
which matters a great deal when your DV is measured in milliseconds.
Configuration knobs (Builder → Tasks → ePAT)
| Setting | Default | Purpose |
|---|---|---|
trials |
20 | Target number of valid trials. |
trial_duration_sec |
30 | Max seconds per trial before auto-cancel. |
retry_budget |
30 | Hard cap on total attempts (valid + failed). |
sqi_threshold |
0.008 | Minimum ePATCore signal-quality value for trial acceptance. |
confidence_ratings |
on | Ask for 1–5 confidence after each trial. |
two_phase_practice |
on | Include tone-to-tone + tone-to-heartbeat practice. |
body_map |
on | Show the post-trial body-location picker. |
ePAT is flagged BETA. The algorithm is stable and the task is usable, but the published-psychometrics validation of the ecological adaptation is still in progress. If you publish ePAT data collected via EMA Forge, please cite both the underlying PAT lineage (Plans et al., 2021; Palmer et al., 2025) and this repository — full citation list in Acknowledgments.
Heartbeat Counting Task (HCT) BETA
The Heartbeat Counting Task is a classic measure of cardiac interoceptive accuracy
(Schandry, 1981). Participants silently count the heartbeats they perceive over a set
of timed intervals. The accuracy score for each interval is
1 - |actual - reported| / actual, where actual is the objective
beat count and reported is the participant's count. Mean accuracy across
intervals is the canonical Schandry score.
EMA Forge's HCT module reuses the same camera-PPG pipeline (ePATCore)
that powers the ePAT and the inline heart-rate question type, so objective beat counts
during each interval are computed with the WABP onset detector and signal-quality
gating described in HR Capture (PPG). Participants hold
their fingertip over the rear camera + flashlight for the duration of the task.
Methodological options (click to expand)
HCT validity has been debated extensively (see Desmedt et al., 2018; Ring & Brener, 2018). The Builder exposes the parameters most likely to matter for that debate:
- Instruction variant. "Count perceived heartbeats" (Schandry's original) versus "Estimate heartbeats" (Brener/Ring). These prompt different cognitive strategies and produce non-equivalent scores; pick deliberately.
- Custom instructions. Free-text override of the variant default, for labs running specific protocols.
- Counting screen visibility. Blank, subtle progress ring, elapsed timer, or both. Hiding the timer is closer to the original Schandry protocol; showing it matches some smartphone-adapted variants.
- Interval set. Comma-separated list. Default is the reduced 25/35/45 s set for EMA contexts; the classic Schandry set is 25, 35, 45, 50, 55, 100. Researcher-supplied sets are honored verbatim.
- Order randomization. Default on; toggle off for fixed-order replications.
- Practice interval. One short practice interval (default 15 s) precedes the real intervals when enabled. Practice trials are recorded but not included in summary statistics.
- Per-interval confidence. 0–10 slider after each interval. Required for Garfinkel et al.'s interoceptive awareness metric (Pearson r between accuracy and confidence across intervals), which the runtime computes when at least three valid (accuracy, confidence) pairs exist.
- Body map. Optional sensation-localization prompt every N intervals, identical to the ePAT body map.
Quality control mirrors ePAT exactly: the same SQI watchdog overlays the same "make the circle red" prompts, and a session-level retry budget silently re-runs intervals that fail signal-quality gating (low SQI for >50% of the interval, or fewer than 5 detected beats).
HCT JSON output shape (click to expand)
{
"type": "hct_response",
"startedAt": "2026-04-30T08:04:11Z",
"baseline": { "recordedHR": [72, 71, ...], "totalBeats": 60, "finalSqi": 0.012, ... },
"intervals": [
{
"intervalIndex": 1, "isPractice": false,
"duration_sec": 35, "durationMs_actual": 35012,
"reportedCount": 32, "actualBeats": 41,
"accuracy": 0.7805,
"confidence": 6, "bodyPos": 1,
"recordedHR": [70, 71, ...], "ibi_series": [810, 795, ...],
"qualitySummary": { "cleanBeats": 39, "flaggedBeats": 2,
"sqiBadSeconds": 0.5, "sqiFinalValue": 0.012, ... }
}, ...
],
"practices": [...],
"summary": {
"valid_intervals": 3, "practice_intervals": 1,
"mean_accuracy": 0.81,
"mean_confidence": 5.7,
"interoceptive_awareness": 0.42,
"mean_sqi": 0.011,
"instruction_variant": "count",
"randomized": true,
"intervals_sec": [25, 35, 45]
}
}
Implicit Association Task (IAT)
The IAT (Greenwald, McGhee & Schwartz, 1998) measures the strength of automatic associations between concept pairs by comparing response times across compatible and incompatible sorting conditions. EMA Forge implements the standard 7-block IAT with D-score scoring (Greenwald, Nosek & Banaji, 2003) computed entirely in the participant's browser at session end. No server, no camera, no hardware requirements beyond a touchscreen.
Because the IAT is RT-based, mobile deployment requires deliberate engineering choices that differ from desktop implementations:
- touchstart, not click. Response time is recorded from
touchstart, which fires at finger contact (~50–100 ms earlier thanclick).clickfires as a fallback for desktop/preview. - rAF-anchored stimulus onset. The stimulus word is written to the
DOM and then
t₀ = performance.now()is captured inside arequestAnimationFramecallback — not at the moment of theinnerHTMLwrite. This ensurest₀reflects pixel-on-screen rather than JS execution time, which can diverge by a full frame (16 ms at 60 Hz) or more under load. - Half-screen tap zones. The left and right halves of the screen are independent invisible buttons. A hairline center divider provides a visual affordance. Tap targets are intentionally the full screen height, not thumb-zone strips, to minimize motor error variance.
D-score algorithm (click to expand)
The implementation follows Greenwald, Nosek & Banaji (2003, JPSP), error replacement strategy B:
- Exclusions. Trials with RT < 300 ms (anticipatory) or > 10,000 ms (disengaged) are flagged as excluded and removed before all subsequent steps.
- Error penalty. For each pairing block set, compute the mean RT of correct trials. Each error trial's RT is replaced by that mean + 600 ms.
- Block means. Mean RT is computed across the error-replaced trial set for blocks 3+4 (pairing 1) and blocks 6+7 (pairing 2).
- Pooled SD. SD is computed across all valid trials from both pairing sets pooled together — not the average of within-block SDs. The same error-replaced RTs are used.
- D = (M₂ − M₁) / SDpooled. Sign is oriented so
positive D = target A + positive attribute faster (compatible). Whether pairing
1 is the compatible condition depends on
block_order_variant(0 or 1), which is logged in the summary and determined byparticipantId % 2.
Fast-responder flag. If more than 10% of trials in any block
have RT < 300 ms, fast_responder: true is set in the
summary. This is a flag, not an auto-exclusion — the convention in the literature
is to report and let the analyst decide.
Block structure (click to expand)
| Block | Content | Default trials | Used for D |
|---|---|---|---|
| 1 | Target A practice | 20 | No |
| 2 | Attribute practice | 20 | No |
| 3 | Combined practice — pairing 1 | 20 | Yes |
| 4 | Combined critical — pairing 1 | 40 | Yes |
| 5 | Target reversal practice | 40 | No |
| 6 | Combined practice — pairing 2 | 20 | Yes |
| 7 | Combined critical — pairing 2 | 40 | Yes |
All trial counts are configurable in the Builder. Blocks 3+4 and 6+7 are the critical blocks used for D-score computation (Greenwald et al., 2003). Practice blocks can be disabled as a group in the Builder, though this is not recommended — participants need to learn the categorization before the combined critical blocks.
Counterbalancing (click to expand)
Which target category pairs with positive attributes in blocks 3/4 is determined
automatically by participantId % 2:
- Variant 0 (even PIDs): Target A + Positive on left in blocks 3/4; Target A + Negative on left in blocks 6/7.
- Variant 1 (odd PIDs): Target A + Negative on left in blocks 3/4; Target A + Positive on left in blocks 6/7.
This gives approximately equal assignment without researcher intervention and
is fully deterministic — you can always reconstruct which order a participant
received from block_order_variant in the output. Include it as a
covariate or between-subjects factor in any model where order effects are a
concern.
IAT JSON output shape (click to expand)
{
"type": "iat_response",
"startedAt": "2026-04-30T08:12:44Z",
"trials": [
{
"block_index": 0,
"block_id": 1,
"block_label": "Practice: Flowers",
"pairing_id": null,
"critical_for_d": false,
"trial_n_in_block": 1,
"trial_n_overall": 1,
"stimulus": "Orchid",
"category": "target_a",
"correct_side": "left",
"response_side": "left",
"correct": true,
"rt_ms": 621.4,
"excluded": false,
"exclude_reason": null,
"timestamp": "2026-04-30T08:12:45Z"
},
...
],
"summary": {
"d_score": 0.482,
"d_mean_pairing1": 712.3,
"d_mean_pairing2": 891.6,
"d_sd_pooled": 372.1,
"d_n_pairing1": 118,
"d_n_pairing2": 117,
"block_order_variant": 0,
"fast_responder": false,
"total_trials": 200,
"excluded_trials": 3,
"exclusion_rate": 0.015,
"target_a_label": "Flowers",
"target_b_label": "Insects",
"attr_pos_label": "Pleasant",
"attr_neg_label": "Unpleasant",
"block_stats": [
{
"block_id": 1,
"n_total": 20, "n_valid": 20, "n_excluded": 0,
"n_errors": 2, "error_rate": 0.1,
"mean_rt": 698.2,
"fast_trial_count": 0, "fast_trial_rate": 0.0
},
...
]
}
}
In R, to get one row per participant with the D-score and block order:
library(jsonlite)
library(dplyr)
library(purrr)
sessions <- list.files("data/", pattern = "\\.json$", full.names = TRUE) |>
map(read_json)
iat_summary <- sessions |>
map_dfr(function(s) {
iat <- keep(s$data, ~ .$type == "iat_response")
if (!length(iat)) return(NULL)
sm <- iat[[1]]$summary
tibble(
pid = s$participantId,
day = s$day,
d_score = sm$d_score,
block_order = sm$block_order_variant,
fast_responder = sm$fast_responder,
exclusion_rate = sm$exclusion_rate,
mean_rt_p1 = sm$d_mean_pairing1,
mean_rt_p2 = sm$d_mean_pairing2
)
})
# Trial-level data (for block-by-block RT analysis)
iat_trials <- sessions |>
map_dfr(function(s) {
iat <- keep(s$data, ~ .$type == "iat_response")
if (!length(iat)) return(NULL)
iat[[1]]$trials |>
map_dfr(as_tibble) |>
mutate(pid = s$participantId, day = s$day)
})
Conditional Tasks
A task step in a window's phase sequence can carry a condition that gates whether it runs at all:
{ kind: "task", id: "epat",
condition: { question_id: "q_hr_1", operator: "gt", value: 80 } }
At runtime, EMA Forge evaluates the condition against the participant's responses collected earlier in the same session. If the condition is false, the step is silently skipped — as if it were never in the sequence. Typical use cases:
- "Only run ePAT if resting HR is elevated" — gate on a prior HR capture's BPM.
- "Only show the post-task stress items if the participant reported feeling stressed beforehand" — gate on a slider threshold.
- "Skip the cognitive task on the 3rd daily window" — gate on window ID.
Privacy Model
Written plainly, because IRBs will ask you to restate this in your protocol:
- Response data transmission is opt-in. By default, data is generated, stored, and downloaded entirely on the participant's device, and you receive it via whatever channel the participant uses to return files. If the researcher configures a webhook (Study tab → Webhook URL), the session JSON is POSTed to that specific endpoint at session end instead. No transmission happens to any endpoint other than the one you configure. The webhook field is off by default.
- The static host sees page-load requests. Your GitHub Pages / Netlify / institutional
host will receive HTTP requests when the participant opens a link. These requests include timestamp, IP,
User-Agent, and the full URL (including
?id=). If participant ID alone is identifiable, choose an unlinkable ID (short opaque strings, not names or MRNs). - No third-party analytics, no CDN fetches, no font-provider calls. The exported study has zero outbound network calls at runtime other than (a) the initial page load from your host, and (b) the single webhook POST at session end if configured. You can verify this in Chrome DevTools → Network.
- Camera / microphone access (HR, ePAT) is handled entirely through the browser's
getUserMediapermission prompt. The media stream never leaves the device; only the derived signal (BPM, IBIs, phase offsets) is stored. - Consent is captured as a boolean plus a timestamp in the onboarding JSON payload. The full consent text as the participant saw it is logged alongside.
IRB Boilerplate
Two variants depending on your data-return design — use the one that matches your protocol.
Variant A: Manual participant return (no webhook)
Participant response data will be collected via a custom web-based application (EMA Forge, an open-source static web tool). The application operates entirely in the participant's mobile browser: response data is generated and stored on the participant's device and is not transmitted to any central server in the course of normal operation. Participants return completed session files to the research team via [INSTITUTIONAL SECURE CHANNEL — e.g., REDCap file upload, institutional SFTP]. No protected health information or direct identifiers are included in the study data; participants are identified only by an opaque study ID. Physiological measurements (heart rate via photoplethysmography) are derived from the participant's smartphone camera locally; raw video is not stored or transmitted. The application source code is open and available for review at github.com/keeganwhitacre/emaforge.
Variant B: Webhook auto-upload to a defined endpoint
Participant response data will be collected via a custom web-based application (EMA Forge, an open-source static web tool). The application operates in the participant's mobile browser and, at the end of each session, transmits the session's response data via HTTPS POST to a single pre-configured endpoint controlled by the research team at [ENDPOINT DESCRIPTION — e.g., an institutionally-hosted endpoint writing to a secured REDCap project / an institutional Google Workspace sheet under the PI's account / an AWS Lambda writing to an encrypted S3 bucket in the institution's AWS organization]. No data is transmitted to any other endpoint. If network transmission fails, the application falls back to on-device storage and manual return via [FALLBACK CHANNEL]. No protected health information or direct identifiers are included in the study data; participants are identified only by an opaque study ID. Physiological measurements (heart rate via photoplethysmography) are derived from the participant's smartphone camera locally; raw video is not stored or transmitted. The application source code is open and available for review at github.com/keeganwhitacre/emaforge.
Known Limitations
- Admin is operational, not a full clinical data platform. The Cloudflare Admin workspace provides live response counts, recent sessions, protected exports, roster/delivery controls, and descriptive analysis; confirmatory analysis and higher-risk access controls remain the researcher's responsibility.
- Missed-prompt detection requires event data. Response files alone cannot reveal prompts that produced no file. Analyze therefore leaves completion unavailable unless a roster and scheduled prompt-event log or a synthetic schedule manifest is present.
- Device heterogeneity. HR/ePAT quality depends on camera sensor, torch brightness, and case thickness. Pilot across a range of devices before locking your protocol.
- Browser storage can be cleared. A participant who "clears website data" mid-study loses their in-progress resume state (but not already-downloaded session files). The completion lock also resets.
- Single-file exports don't let you patch typos. Use the static-hosting bundle if you anticipate edits after deployment.
- Accessibility is a work in progress. Screen-reader support for sliders and the affect grid is incomplete. Audit against your accessibility requirements before broad enrollment.
Roadmap
Auto-upload with Webhooks BETA
The recommended prepared Cloudflare study automatically posts each completed session to the deployment's
same-host /submit endpoint. "Save Local Copy" is the safety fallback when storage is unavailable.
Manual return remains available when no receiver is configured.
For independent hosting, the webhook option provides the same automation. When you paste an
approved HTTPS URL into the Receiver field, the exported study silently POSTs the full session JSON to that URL at
save-time. The participant sees "Data uploaded successfully" only after the receiver responds with readable
JSON containing {"status":"success"}. A failed or unreadable acknowledgment offers a local
download. Confirm storage on the receiver before enrollment; an HTTP 200 alone does not prove data was saved.
Design tradeoff worth understanding: enabling a webhook means you are now transmitting response data to an endpoint of your choosing. The core "serverless" guarantee (data stays on-device) is opt-in to preserve. If your IRB protocol specifies manual return, leave the field blank. If your IRB approves direct electronic return to a specific institutional endpoint, webhooks make the data pipeline hands-off for participants.
What gets sent
The runtime POSTs a JSON body to your endpoint with this shape:
{
"submission_id": "s_a1b2c3",
"participant_id": "104",
"day": 2,
"window_id": "morning",
"session_data": {
"participantId": "104",
"sessionId": "s_a1b2c3",
"day": 2,
"type": "ema_with_task",
"status": "complete",
"startedAt": "2026-04-22T08:03:11Z",
"completedAt": "2026-04-22T08:06:48Z",
"data": [ /* ema_response, epat_response, etc. — full payload */ ]
}
}
The top fields include a stable submission_id for receiver deduplication and convenient
dispatch/filtering on your
receiving end. The session_data object is the same JSON documented in JSON Schema.
Content-Type is text/plain, not application/json. This is
deliberate — it's the standard workaround that keeps the browser from firing a CORS preflight (OPTIONS)
request, which Google Apps Script and many other simple receivers cannot answer. Your endpoint must
parse the body as JSON despite the text/plain header. Every code example below does this.
Webhook recipes by receiver
These recipes are examples that require testing in your own hosting and institutional environment. The endpoint
must return readable JSON with {"status":"success"} after durable storage and permit the study origin
through CORS. The Google Apps Script option may save a row even when a browser cannot read its redirected response.
EMA Forge's webhook is a plain HTTPS POST with a JSON body. Anything that
can accept that and return a readable success receipt after saving can work. The table below summarises the
recipes documented in this section; pick the one whose tradeoffs match your
IRB, your IT environment, and your willingness to do setup work.
| Receiver | Cost | Setup | When to pick it |
|---|---|---|---|
| Google Apps Script → Sheets | Free | ~5 min | Illustrative only: browser acknowledgment across Apps Script redirects is unverified. |
| Cloudflare Worker → R2 | Free tiers available; billing setup required for R2 | Guided deploy; allow time for Cloudflare account and R2 setup | Recommended reference path. Private object storage, verified receipts, duplicate protection. |
| Azure Function → SharePoint | Free (consumption tier) | ~30 min + IT signoff | Institutional Microsoft 365 mandate. Strongest IRB story for university-managed tenants. |
| Power Automate → Excel | Premium licence required | ~10 min | You already have Power Automate Premium. Simpler than the Azure path if you do. |
| PHP proxy → REDCap | Free (uses existing lab server) | ~30 min + lab IT | IRB requires data stay on institutional infrastructure end-to-end. |
| AWS Lambda → S3 | ≈$0–3/month | ~45 min | Lab is standardised on AWS, or you want presigned-URL access patterns. |
Google Sheets via Apps Script BETA
An example receiver for testing with non-sensitive data. Apps Script Content Service redirects its response; browser cross-origin access and institutional approval must be checked before relying on it for participant data.
Full step-by-step setup (click to expand)
- Open sheets.google.com and create a new blank spreadsheet. Name it whatever you want your study data labelled as.
- From the menu bar, Extensions → Apps Script.
- Delete the placeholder
function myFunction() {}and paste the script below. - Save. Name the project anything.
- Deploy → New deployment. Choose type Web app.
- Set Execute as: Me, Who has access: Anyone. Necessary because the exported study has no Google credentials. Your endpoint URL is effectively the credential — treat it as a secret.
- Click Deploy. Authorize when prompted. Copy the Web app URL
(starts with
https://script.google.com/macros/s/...). - Paste into EMA Forge's Builder → Study tab → Webhook URL field.
- Re-export. Test end-to-end with a dummy participant ID.
function doPost(e) {
try {
const data = JSON.parse(e.postData.contents);
const sheet = SpreadsheetApp.getActiveSpreadsheet().getActiveSheet();
if (sheet.getLastRow() === 0) {
sheet.appendRow(['Timestamp', 'Participant ID', 'Day', 'Window', 'Raw JSON']);
sheet.getRange(1, 1, 1, 5).setFontWeight("bold");
sheet.setFrozenRows(1);
}
sheet.appendRow([
new Date(),
data.participant_id || 'Unknown',
data.day || 'Unknown',
data.window_id || 'Unknown',
JSON.stringify(data.session_data || data)
]);
return ContentService
.createTextOutput(JSON.stringify({status: 'success', submission_id: data.submission_id}))
.setMimeType(ContentService.MimeType.JSON);
} catch (error) {
return ContentService
.createTextOutput(JSON.stringify({error: error.message}))
.setMimeType(ContentService.MimeType.JSON);
}
}
When editing, redeploy via Manage deployments → Edit → New version (keeps the same URL).
Apps Script rate limits are generous (~1000 sessions/day before issues). Beyond that, switch to Cloudflare or Azure.
Cloudflare study host → private R2 BETA
This is EMA Forge's recommended path for hosting, study administration, optional messaging, and automatic data return. One Worker serves the participant study and stores one JSON object per session, uses the stable submission ID as its key, confirms a successful write before acknowledging the participant app, treats an exact retry as a duplicate, and refuses changed data that attempts to reuse an existing ID. The private bucket has no public read route. Study installation, roster management, monitoring, browser-local analysis, and exports require the deployment's admin token.
One deployment per study. A Worker binds one private R2 bucket and maintains one active study, roster, and response collection. For a new study in the same Cloudflare account, create another Worker with a unique name; the new template leaves the STUDY_DATA bucket name unset so Cloudflare provisions a separate bucket automatically. Confirm that the new Worker's binding points to its own bucket before installation. Older deployments made with a fixed default bucket name may still need a one-time binding correction. If Admin reports a storage conflict, connect a new empty bucket before installing; never clear the existing study's bucket. Reinstalling a protocol for the current study preserves historical responses; installing a differently named study is blocked so those records are not mixed.
Updating an existing study: Download the new prepared HTML and install it in the existing Study Admin. Keep the existing Worker, URL, and R2 bucket. If an EMA Forge feature also changes Worker code, choose Update an existing Worker’s code → Download Worker update in Review & Deploy. Select the existing Worker in Cloudflare Workers & Pages, choose Edit code, replace its code with that self-contained file, and deploy while preserving its STUDY_DATA binding and admin secret. The raw worker.mjs has a separate import and is not the single-file update. The Create new study host button is for a separate study, not a software update. If running several independent studies, each has its own Worker and bucket; Worker code changes must currently be applied to each. Do not delete a Worker to update it.
In Review & Deploy, download the prepared Cloudflare study and open Deploy to Cloudflare.
Cloudflare clones the public template, provisions and binds private R2 storage, and shows a masked example
ADMIN_TOKEN. Replace that example with your own 32+ character password; EMA Forge cannot recover it.
Twilio is optional and can be connected from the protected Study Admin after deployment. The template enables its workers.dev route automatically.
Copy the address Cloudflare shows and paste it into EMA Forge under Hosted study URL; a bare hostname
or full HTTPS URL both work. Select Open admin. A new deployment opens directly on
Install study; upload the downloaded file. EMA Forge never receives
the researcher's Cloudflare credentials, token, study file, or participant data.
The hosted URL is saved with the local builder workspace, so returning to Review & Deploy restores clearly labeled participant, administration, and connection-check destinations without another search through Cloudflare.
Cost: EMA Forge is free. A modest study may stay within Workers Free limits and R2 Standard free allowances. R2 still requires account activation/checkout, and overages can cost money. Twilio SMS, sending numbers, and carrier fees are separate. Check expected usage and retention before recruiting.
IRB and institutional review: Researcher-owned Cloudflare storage makes the data flow inspectable, but it does not establish approval or HIPAA compliance. Document the provider, data types, consent process, access controls, region, retention and deletion, and incident response for your IRB and IT/security review. If your institution requires an approved vendor or institutional storage, use that pathway instead.
If the admin page reports Unauthorized, do not recreate the deployment. Open
Worker → Settings → Variables and Secrets, add or replace a runtime Secret named exactly
ADMIN_TOKEN, use a value of at least 32 characters, and deploy the settings change. Do not place the token
value in wrangler.jsonc or commit it to Git.
Run /check.html after installation. It presents a clear success or failure instead of raw JSON.
Synthetic checks are stored under setup-tests/ and are excluded from response counts, Analyze, and
response exports. Real sessions are stored under sessions/.
Independent receiver setup (click to expand)
- In Review & Deploy, enter the eventual hosted study URL and select Download receiver starter.
- Extract the zip. It includes the receiver, configuration, and a short README.
- Create a Cloudflare account and enable R2. From the extracted folder, run
npx wrangler login. - Run
npx wrangler r2 bucket create ema-forge-study-data. Keep the bucket private. - Open
wrangler.jsonc. ConfirmALLOWED_ORIGINSexactly matches the hosted study's origin. Change the Worker and bucket names when deploying a separate study. - Run
npx wrangler deploy. Copy the HTTPS URL and append/submit. - Paste that complete URL into the Builder's Receiver URL field, export the static bundle, and upload every file.
- Open the hosted
check.htmland send a synthetic test. Confirm the displayed ID exists undersetup-tests/in R2. Then complete a real test session and inspect its object undersessions/.
Security boundary: the allowed-origin check controls ordinary browser access; it is not authentication because a non-browser client can set an Origin header. The receiver URL is visible in the study bundle and cannot be a secret. Use opaque participant IDs and assess consent, abuse controls, retention, and institutional approval for the sensitivity of the study. Deploy a separate receiver and bucket per study.
To pull data for analysis, R2 speaks S3 API. From R:
library(aws.s3)
Sys.setenv(
AWS_ACCESS_KEY_ID = "<your R2 access key>",
AWS_SECRET_ACCESS_KEY = "<your R2 secret key>",
AWS_S3_ENDPOINT = "<account_id>.r2.cloudflarestorage.com"
)
# Sync entire bucket to local
files <- get_bucket_df(bucket = "ema-forge-sessions", region = "auto")
for (key in files$Key) {
save_object(key, bucket = "ema-forge-sessions",
file = file.path("data", basename(key)))
}
Generate R2 API tokens at Cloudflare dashboard → R2 → Manage R2 API Tokens.
Azure Function → SharePoint List STABLE
The institutional Microsoft path. Best fit when your university mandates Microsoft 365 for research data storage, when your IRB prefers data stay inside the institutional tenant, or when your IT contact will resist approving non-Microsoft cloud services. Azure Functions on the consumption plan are free up to 1 million executions per month, which is well above any plausible EMA workload.
This recipe requires a one-time app registration in Azure AD, which
typically means asking your central IT to grant the
Sites.ReadWrite.All application permission. Most university
IT shops process this routinely.
Full step-by-step setup (click to expand)
One-time Azure AD setup
- Azure Portal → App Registrations → New registration.
Name:
EMA-Forge-Receiver. Account type: single tenant. Click Register. - From the new app's overview, copy Application (client) ID and Directory (tenant) ID.
- Certificates & secrets → New client secret. Copy the secret Value immediately — it's never shown again.
- API permissions → Add a permission → Microsoft Graph → Application permissions → Sites.ReadWrite.All. Add. Then Grant admin consent (your IT will need to do this click; the rest you can do yourself).
SharePoint list setup
- In your team's SharePoint site, create a new List
called
EMA Sessions. - Add columns:
Participant— Single line of textDay— NumberWindow— Single line of textSessionJSON— Multiple lines of text, plain text, increase limit to 250,000 charactersReceivedAt— Date and Time
- Get the site ID and list ID via Graph Explorer:
- Site ID:
GET https://graph.microsoft.com/v1.0/sites/<tenant>.sharepoint.com:/sites/<site-name> - List ID:
GET https://graph.microsoft.com/v1.0/sites/<site-id>/lists— find your list by displayName
- Site ID:
Azure Function deployment
- Install Azure Functions Core Tools and the Azure CLI.
- Create a new function project:
func init ema-receiver --javascript --model V4 cd ema-receiver func new --name SaveSession --template "HTTP trigger" --authlevel "function" - Replace
src/functions/SaveSession.jswith:
const { app } = require('@azure/functions');
const { Client } = require('@microsoft/microsoft-graph-client');
const { ClientSecretCredential } = require('@azure/identity');
require('isomorphic-fetch');
app.http('SaveSession', {
methods: ['POST', 'OPTIONS'],
authLevel: 'function',
handler: async (req, ctx) => {
const cors = {
'Access-Control-Allow-Origin': '*',
'Access-Control-Allow-Methods': 'POST, OPTIONS',
'Access-Control-Allow-Headers': 'Content-Type',
};
if (req.method === 'OPTIONS') return { status: 200, headers: cors };
try {
const raw = await req.text();
const data = JSON.parse(raw);
const credential = new ClientSecretCredential(
process.env.TENANT_ID,
process.env.CLIENT_ID,
process.env.CLIENT_SECRET
);
const graph = Client.initWithMiddleware({
authProvider: {
getAccessToken: async () =>
(await credential.getToken('https://graph.microsoft.com/.default')).token
}
});
const siteId = process.env.SHAREPOINT_SITE_ID;
const listId = process.env.SHAREPOINT_LIST_ID;
await graph
.api(`/sites/${siteId}/lists/${listId}/items`)
.post({
fields: {
Title: `${data.participant_id}_D${data.day}_${data.window_id}`,
Participant: String(data.participant_id || ''),
Day: Number(data.day) || 0,
Window: String(data.window_id || ''),
ReceivedAt: new Date().toISOString(),
SessionJSON: JSON.stringify(data.session_data || {}).slice(0, 250000)
}
});
return {
status: 200,
headers: { ...cors, 'Content-Type': 'application/json' },
body: JSON.stringify({ status: 'success' })
};
} catch (err) {
ctx.error(err);
return {
status: 500,
headers: { ...cors, 'Content-Type': 'application/json' },
body: JSON.stringify({ error: err.message })
};
}
}
});
npm install @azure/functions @azure/identity @microsoft/microsoft-graph-client isomorphic-fetch
- Deploy:
az login az functionapp create --resource-group <rg> --consumption-plan-location <region> \ --runtime node --runtime-version 20 --functions-version 4 \ --name ema-receiver-<you> --storage-account <storage> func azure functionapp publish ema-receiver-<you> - Set environment variables in the Function App Configuration:
TENANT_ID,CLIENT_ID,CLIENT_SECRET,SHAREPOINT_SITE_ID,SHAREPOINT_LIST_ID. - Get the function URL (includes a
?code=key acting as a shared secret): Azure Portal → your Function → Functions → SaveSession → Get Function Url. - Paste into EMA Forge's Webhook URL field. Re-export.
For Ohio State users specifically
OSU permits use of the university-configured Azure tenant for non-medical-center research data classified at S3 (Private) or below. Behavioural EMA data with coded participant IDs and no PHI is S3. The OSU Azure tenant lives under the university's Microsoft enterprise agreement and meets OSU's Information Security Standard for S3 data without requiring a separate Information Security Risk Assessment. Email your college Security Coordinator before submitting your IRB protocol to confirm the setup matches their expectations for your specific data.
Power Automate → Excel STABLE
Simpler than the Azure Function path, but the "When an HTTP request is received" trigger is a Premium Power Automate connector that most institutional Microsoft 365 plans do not include. Before investing time, check with IT whether your Microsoft 365 plan has Premium Power Automate. If not, use the Azure Function path instead.
Full step-by-step setup (click to expand)
Prepare the Excel destination
- Create a new Excel file in OneDrive for Business:
EMA-Forge-Data.xlsx. - Add headers in row 1:
Timestamp | Participant_ID | Day | Window | Raw_JSON. - Select the range and press Ctrl+T to format as a table.
Name the table
SessionData. Power Automate's Excel connector cannot see unformatted ranges — the table step is required.
Build the flow
- Go to make.powerautomate.com. New flow → Instant cloud flow → "When an HTTP request is received".
- In the trigger, paste this Request Body JSON Schema:
{ "type": "object", "properties": { "participant_id": {"type": "string"}, "day": {"type": "integer"}, "window_id": {"type": "string"}, "session_data": {"type": "object"} } } - Add action: Compose. Inputs:
@{string(triggerBody()?['session_data'])}. This stringifies the nested session data so Excel can store it as a single cell. - Add action: Excel Online (Business) → Add a row into a table.
- Location: OneDrive for Business
- Document Library: OneDrive
- File:
/EMA-Forge-Data.xlsx - Table:
SessionData - Row fields:
- Timestamp:
@{utcNow()} - Participant_ID:
@{triggerBody()?['participant_id']} - Day:
@{triggerBody()?['day']} - Window:
@{triggerBody()?['window_id']} - Raw_JSON:
@{outputs('Compose')}
- Timestamp:
- Add action: Response. Status code: 200. Body:
{"status": "success"}. Must be the final action. - Save the flow. Click back into the trigger card — the HTTP POST URL field now contains the endpoint URL. Copy it.
- Paste into EMA Forge's Webhook URL field. Re-export.
Performance caveat: Power Automate must return the Response within 120 seconds or the flow runs asynchronously and the runtime can't tell whether the save succeeded. For standard EMA payloads this is never a problem; for sessions with raw PPG samples (~MB-scale), monitor for timeouts during pilot.
REDCap via institutional PHP proxy STABLE
Direct browser-to-REDCap POSTs are blocked by CORS on most REDCap installations. The realistic path is a one-file PHP proxy hosted on your lab's institutional web server, which accepts the POST from the participant's phone and forwards it to REDCap's API using your project token. This is the path most IRBs prefer because the data path is institutional end-to-end — the participant's phone hits a university domain, which forwards to a university REDCap instance, with no external cloud services involved.
Full step-by-step setup (click to expand)
REDCap project setup
- In your REDCap project, enable the API via Project Setup → API.
- User Rights → assign yourself API rights → generate an API token. Copy it — you'll need it in the PHP file. Treat as a credential.
- Online Designer → create a repeatable instrument called
session. Add fields:session_day— Text Boxsession_window— Text Boxsession_json— Notes Box (long text — verify your REDCap install allows >30k chars per field)session_received— Text Box with date/time validation
PHP proxy
Drop this on your lab's Apache/nginx server. Most psych departments have one — ask your IT contact. The file is dependency-free vanilla PHP and works on any version ≥7.4.
<?php
// ema-forge-redcap-proxy.php
// Forwards EMA Forge sessions into a REDCap project's repeatable instrument.
header('Access-Control-Allow-Origin: *');
header('Access-Control-Allow-Methods: POST, OPTIONS');
header('Access-Control-Allow-Headers: Content-Type');
header('Content-Type: application/json');
if ($_SERVER['REQUEST_METHOD'] === 'OPTIONS') {
http_response_code(200);
exit;
}
$REDCAP_URL = getenv('REDCAP_URL') ?: 'https://redcap.your-uni.edu/api/';
$REDCAP_TOKEN = getenv('REDCAP_TOKEN') ?: 'PASTE_YOUR_TOKEN_HERE';
$SHARED_SECRET = getenv('EMA_SECRET') ?: ''; // optional, recommended
if ($SHARED_SECRET && ($_GET['auth'] ?? '') !== $SHARED_SECRET) {
http_response_code(401);
echo json_encode(['error' => 'unauthorized']);
exit;
}
$body = file_get_contents('php://input');
$data = json_decode($body, true);
if (!$data) {
http_response_code(400);
echo json_encode(['error' => 'invalid JSON']);
exit;
}
$record = [[
'record_id' => $data['participant_id'] ?? 'unknown',
'redcap_event_name' => 'ema_arm_1', // edit to match your project
'redcap_repeat_instrument' => 'session',
'redcap_repeat_instance' => 'new', // REDCap auto-assigns
'session_day' => $data['day'] ?? '',
'session_window' => $data['window_id'] ?? '',
'session_json' => json_encode($data['session_data'] ?? []),
'session_received' => date('Y-m-d H:i:s'),
]];
$payload = http_build_query([
'token' => $REDCAP_TOKEN,
'content' => 'record',
'format' => 'json',
'type' => 'flat',
'data' => json_encode($record),
'overwriteBehavior' => 'normal',
'forceAutoNumber' => 'false',
'returnContent' => 'count',
]);
$ch = curl_init($REDCAP_URL);
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_POSTFIELDS => $payload,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_SSL_VERIFYPEER => true,
CURLOPT_TIMEOUT => 30,
]);
$result = curl_exec($ch);
$code = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
if ($code >= 200 && $code < 300) {
http_response_code(200);
echo json_encode(['status' => 'success']);
} else {
http_response_code(502);
echo json_encode(['error' => 'redcap rejected', 'code' => $code]);
}
- Have lab IT place the file behind HTTPS at e.g.
https://yourlab.dept.osu.edu/ema-proxy.php. - Set environment variables (Apache:
SetEnvin the vhost config; nginx + PHP-FPM:env[REDCAP_TOKEN] = ...in pool config). - Paste
https://yourlab.dept.osu.edu/ema-proxy.php?auth=<shared-secret>into EMA Forge's Webhook URL field. - Re-export. Test end-to-end.
Edit the redcap_event_name to match your project's
arm/event scheme — REDCap will reject the record otherwise.
AWS Lambda → S3 STABLE
For labs standardised on AWS. Cost is roughly $0–3 per month at EMA scale: Lambda's free tier covers ~1M requests/month indefinitely, and S3 storage runs about $0.023 per GB-month. Output shape mirrors the Cloudflare R2 recipe (one JSON file per session, partitioned by participant/day), so analysis pipelines are interchangeable.
Full step-by-step setup (click to expand)
- Create an S3 bucket:
aws s3 mb s3://ema-forge-sessions-<you>. - Block all public access on the bucket (default is correct).
- Create a Lambda function (Node.js 20.x runtime) with this code:
import { S3Client, PutObjectCommand } from '@aws-sdk/client-s3';
const s3 = new S3Client({ region: process.env.AWS_REGION });
const BUCKET = process.env.BUCKET_NAME;
export const handler = async (event) => {
const cors = {
'Access-Control-Allow-Origin': '*',
'Access-Control-Allow-Methods': 'POST, OPTIONS',
'Access-Control-Allow-Headers': 'Content-Type',
};
if (event.requestContext?.http?.method === 'OPTIONS') {
return { statusCode: 200, headers: cors };
}
try {
const body = JSON.parse(event.body);
const pid = String(body.participant_id || 'unknown');
const day = String(body.day || '0');
const win = String(body.window_id || 'na');
const ts = Date.now();
const key = `sessions/${pid}/day-${day}/${win}_${ts}.json`;
await s3.send(new PutObjectCommand({
Bucket: BUCKET,
Key: key,
Body: JSON.stringify(body),
ContentType: 'application/json',
Metadata: { participantid: pid, day, windowid: win }
}));
return {
statusCode: 200,
headers: { ...cors, 'Content-Type': 'application/json' },
body: JSON.stringify({ status: 'success', key })
};
} catch (err) {
return {
statusCode: 500,
headers: { ...cors, 'Content-Type': 'application/json' },
body: JSON.stringify({ error: err.message })
};
}
};
- Lambda execution role: attach an inline policy allowing
s3:PutObjectonarn:aws:s3:::ema-forge-sessions-<you>/*. - Environment variables:
BUCKET_NAME. - Add an API Gateway HTTP API trigger (not REST API — HTTP API is
cheaper and simpler). Configure a default route
POST /. - Enable CORS on the API: allow origin
*, methodsPOST, OPTIONS, headersContent-Type. - Copy the invoke URL. Paste into EMA Forge's Webhook URL field.
For analysis, sync via aws-cli or
aws.s3::save_object() in R. The bucket layout
(sessions/<pid>/day-<n>/<window>_<ts>.json)
makes per-participant or per-day pulls trivial.
Cautions and gotchas
- Webhook URLs in a public study are public. Anyone can inspect exported HTML and POST synthetic data, including if the URL contains a query-string secret. Use server-side validation, study-specific limits and reconciliation; a static webpage cannot keep a shared secret.
- Google Apps Script has quotas and redirects. Capacity depends on the account and payload; verify browser-to-receiver acknowledgment and storage under expected load.
- Apps Script may store a row while the browser cannot read its response. Cross-origin redirects and CORS can cause the participant to see "Upload not confirmed" after the server wrote data. Check the sheet before accepting a manually returned copy to avoid duplicates.
- If you change the webhook URL mid-study, participants who cached the old exported HTML will continue to post to the old URL until they clear site data or you push a new export. Make webhook changes at export boundaries, not mid-enrollment.
- Webhooks do not eliminate the manual-download safety net. By design. If connectivity fails at session end, the "Save Local Copy" button reappears and the participant can still return the file by your IRB-approved channel.
Twilio Integration BETA
Cloudflare-native and optional. EMA Forge can schedule participant-specific SMS from the same Worker that hosts the study and stores responses. The delivery pipeline has automated tests, but it still requires a live test of scheduling, carrier delivery, status callbacks, and STOP handling before enrollment.
Twilio credentials can be entered on the protected /admin page after deployment. They are encrypted in private study storage using the admin token; Cloudflare Worker secrets remain an alternative. Google Sheets
and Apps Script are no longer part of the recommended path. The Worker reads the installed EMA Forge
schedule, evaluates it in each participant's IANA timezone every five minutes, sends an individualized
link, and stores an auditable dispatch record separately from response data.
Setup
- Complete the normal Cloudflare deployment and install the prepared study at
/admin. - Open Install study in Study Admin. Enter your Account SID, Auth Token, and either a sending number or Messaging Service SID, then choose Connect Twilio. The credentials are not displayed again. If the admin token changes, reconnect Twilio.
- In Twilio, set the incoming-message webhook to
https://<your-study-host>/twilio/incomingusing POST. The Worker validates Twilio's signed request before changing a roster record. - Open Participants & delivery in Study Admin. Upload a CSV or add participants, send a test message, and only then enable scheduled messaging.
Roster schema
| Field | Meaning |
|---|---|
participant_id | Unique opaque ID used in the participant link and response envelope. |
phone | E.164 phone number, such as +15551234567. |
timezone | IANA timezone, such as America/New_York. This preserves daylight-saving behavior. |
start_date | Calendar start date in YYYY-MM-DD format. |
status | active, inactive, or opted_out. |
schedule_preferences_json | Optional onboarding preferences containing days and per-window time overrides. |
How delivery works
- A Cloudflare Cron Trigger invokes the Worker's scheduled handler every five minutes.
- The Worker determines the participant's calendar study day and local time, intersects participant preferences with protocol weekdays, and derives a stable randomized minute inside each window.
- A prompt is claimed with a participant/day/window key before the Twilio request. This prevents overlapping scheduler runs from sending the same prompt twice.
- The Worker creates an opaque
/j/…link. Its private R2 mapping contains the participant, day, session, and send time needed to enforce the configured response window and grace period. The SMS does not expose those routing parameters. - Twilio API acceptance, sent, delivered, failed, and opt-out events are retained as separate states. Signed delivery callbacks update the existing dispatch record.
Explicit send failures retry up to three times, no more frequently than every fifteen minutes. An API acceptance is not labeled handset delivery. The Analyze tab reports exactly the states received from Twilio and does not infer delivery from a successful API request.
Admin and analysis
The protected Admin workspace provides the participant link, connection check, live response counts, roster, messaging controls, revocable participant invite links, response summaries, physiology-session counts, recent submissions, raw NDJSON, and the dispatch audit log. Invite links are study-scoped rather than a general-purpose URL shortener. Operational analysis runs in the browser against protected Worker endpoints; participant data are not uploaded back to EMA Forge.
Known limitations
- Twilio setup is still external. Researchers must obtain a Twilio sender and satisfy the applicable messaging registration requirements. Credentials can then be entered in Study Admin.
- Five-minute scheduling is approximate. It is appropriate for EMA windows, not exact-to-the-second interventions.
- Phone numbers are sensitive. They are kept in the private roster, separate from response files, but remain the researcher's responsibility under the approved data-management plan.
- Admin summaries are descriptive. Download the lossless response and dispatch files for confirmatory analysis and archived reproducibility.
- Test before enrollment. Verify at least two timezones, STOP, delivery callbacks, link expiry, and one complete response on the actual participant devices.
Under the Hood
You do not need any of this to run a study. It's here for the small subset of users who want to audit the code, contribute a new task module, or fork the project for a deeply custom use case.
Repository layout
The repository is split into three concerns: (1) what the researcher uses (Builder + Dashboard), (2)
what gets compiled into the participant's study (templates/), and (3) shared styling.
ema-forge/ ├── index.html # Landing page ├── builder.html # Study authoring environment ├── dashboard.html # Local analysis dashboard ├── readme.html # This file ├── ppgtester.html # Standalone PPG pipeline tester │ ├── css/ │ ├── studio.css # Shared design tokens + component styles │ ├── builder.css # Builder workspace layout │ └── dashboard.css # Analyze workspace layout │ ├── library/ # Versioned protocols, packs, and task presets │ ├── js/ │ ├── state.js # Central state + schema (SCHEMA_VERSION) │ ├── storage.js # LocalStorage + project import/export + migrations │ ├── export.js # Compiles study to HTML / zip bundle │ ├── preview.js # Live iframe preview │ │ │ ├── tabs/ # Builder tab controllers │ │ ├── study.js │ │ ├── onboarding.js │ │ ├── questions.js │ │ ├── schedule.js # Windows + phase_sequence editor │ │ ├── tasks.js # Pluggable module registry │ │ └── deployment.js # Routing CSV generator │ │ │ └── dashboard/ │ ├── parser.js # Ingests participant JSON folder │ ├── simulator.js # Deterministic synthetic protocol runs │ └── dashboard.js # Charts, filters, CSV export │ └── templates/ # Source files stitched into the exported study ├── epat-core.js # PPG pipeline + BeatDetector ├── study-base.js # Runtime skeleton ├── module-onboarding.js ├── module-ema.js # EMA block + inline HR capture └── module-epat.js # ePAT task
Key boundary: js/ powers the Builder/Dashboard
(what the researcher interacts with). templates/ is the code that gets stitched into the
participant's study by js/export.js. Never edit templates hoping to change Builder
behavior, or vice versa.
Adding a new task module
The task registry in js/tabs/tasks.js exposes a SETTINGS_RENDERERS object.
Each entry is { html(mod), bind(card, mod) } — html() returns the
settings-panel HTML for the Builder, bind() attaches event listeners. To register a module:
- Add an entry to
state.modulesinjs/state.jswithid,label,desc,enabled, and asettingsobject. - Add a matching renderer in
SETTINGS_RENDERERSinjs/tabs/tasks.js. - Create a runtime module file in
templates/module-<id>.jsthat defines a lifecycle (start,teardown,getData) and push it into the export pipeline injs/export.js.
ePAT is the reference implementation; use it as a template.
Self-hosting the Builder (not recommended)
The Builder is hosted at emaforge.keeganwhitacre.com and that is the recommended way to use it. The hosted version is kept on the current schema and is always up to date; self-hosting introduces version-skew risk (see Schema Versioning) that can corrupt saved projects across researchers working on the same study.
That said, if you need to self-host — institutional policy forbids external tools, you want to run an older schema version indefinitely, or you're developing against the code — clone the repo and serve it with any static-file server:
git clone https://github.com/keeganwhitacre/emaforge.git
cd emaforge
python -m http.server 8000
# then open http://localhost:8000
Camera access requires HTTPS, so for HR/ePAT work you'll need a real certificate — file://
and plain HTTP over non-localhost will both fail.
Credit & Citation
EMA Forge was created by Keegan Whitacre at the Affective Science Lab, The Ohio State University. If you use EMA Forge in a published study, please cite:
Whitacre, K. (2026). EMA Forge: A serverless toolkit for
ecological momentary assessment and digital phenotyping.
Affective Science Lab, The Ohio State University.
https://github.com/keeganwhitacre/emaforge
For ePAT-specific methods, also describe the implementation (PPG pipeline + AudioContext-scheduled stimulus + dial-alignment response) in your Methods section. Issues, pull requests, and replication reports are welcome.
License
EMA Forge is released under the MIT License — free for academic and commercial use, modification, and
redistribution. See LICENSE in the repository for the full text.
Acknowledgments & Prior Work
EMA Forge stands on a substantial body of prior work. The ePAT task in particular is not a novel invention but an ecological adaptation of an established psychophysical paradigm. Proper attribution here matters both scholarly and practically — if you publish using these modules, these are the citations your Methods section owes.
The Phase Adjustment Task (PAT) lineage
The core psychophysical logic of the ePAT — aligning an auditory tone to felt heartbeat via a continuous-adjustment response — is adapted from the Phase Adjustment Task developed by Plans, Ponzo, Morelli, Cairo, Ring, Keating, Cunningham, Catmur, Murphy & Bird, with subsequent refinements in PAT 2.0.
- Plans, D., Ponzo, S., Morelli, D., Cairo, M., Ring, C., Keating, C. T., Cunningham, A. C., Catmur, C., Murphy, J., & Bird, G. (2021). Measuring interoception: The phase adjustment task. Biological Psychology, 165, 108171.
- Palmer, C., Murphy, J., Bird, G., et al. (2025). Refinements of the Phase Adjustment Task (PAT 2.0). Preprint. doi:10.31219/osf.io/4qtwv.
- Original reference implementation (Swift/iOS): huma-engineering/Phase-Adjustment-Task.
The ePAT's contribution is specifically the ecological framing: porting the paradigm to a participant-owned smartphone, using camera PPG rather than a dedicated pulse oximeter, scheduling it within an EMA protocol rather than as a discrete lab session, and treating the resulting phase-offset trajectory as a time-varying individual-difference signal rather than a single-point measurement.
Beat detection — WABP
The PPG peak-detection routine is a JavaScript port of the WABP (Waveform Analysis for Blood Pressure) algorithm, originally developed for arterial blood pressure onset detection and released on PhysioNet. The algorithm generalizes well from arterial pressure to PPG because both signals share the characteristic upstroke-dominant morphology the algorithm was designed around.
- Zong, W., Heldt, T., Moody, G. B., & Mark, R. G. (2003). An open-source algorithm to detect onset of arterial blood pressure pulses. Computers in Cardiology, 30, 259–262.
- Reference C implementation: PhysioNet
wabp.c.
Camera-based PPG acquisition
The browser-side camera acquisition approach — sampling the red channel of a rear-camera video stream under torch illumination — was informed by Richard Moore's open-source demonstrator, which established the feasibility of the pattern in a pure web environment.
- Moore, R. heart-rate-monitor. github.com/richrd/heart-rate-monitor.
Conceptual framework
The decision to operationalize interoceptive accuracy as an ecological, time-varying construct — and to embed it within a broader affective-science EMA protocol — is grounded in the Theory of Constructed Emotion and contemporary work on interoceptive predictive processing. The ePAT is one instrument within that larger program; it is not itself a complete theory.
If you use the ePAT in a published study, please cite the PAT and PAT 2.0 papers alongside EMA Forge. The ePAT is an implementation and ecological adaptation, not an independent paradigm — its validity inherits from that lineage.
Maintained by Keegan Whitacre, OSU Affective Science Lab.