STOA Scenario Pack

AI INTEGRATION SCENARIO.

Where should AI mine data, draft cost estimates, surface feasibility issues, and support employee interactions while keeping human validation in the loop?

Assist — AI suggests, human confirms
Automate — AI acts, notifies after
Escalate — AI surfaces, human resolves
5
Actors
5
Triggers
39
Cards
84
Quality Score
↓ Scroll to explore
01 · The Users

The Users

Five actors interact with or depend on the AI layer. The workshop identified both back-office mining agents and employee-facing interaction as important parts of the MVP conversation.

Actor · Gallagher
HR / Mobility Manager
The platform's primary daily user. Manages assignment populations, runs compliance workflows, and approves AI-generated recommendations. Needs AI to surface next best actions, pre-fill repetitive data, and flag approaching deadlines — without AI making final calls that carry legal weight.
Actor · Employee
Relocating Employee
Receives AI-curated communications, personalised relocation checklists, and proactive reminders about visa appointments, housing deadlines, and tax filing windows. Must feel guided, not overwhelmed — AI must know when to inform and when to prompt action.
Actor · Gallagher
Compliance Officer
Relies on AI to surface exceptions, jurisdiction-specific rule changes, and edge cases that fall outside standard policy. Does not want to manually scan every case — AI must escalate only high-confidence anomalies with supporting context attached.
Actor · Platform
GMS Platform AI Engine
The AI agent layer embedded in the platform. Reads assignment state, policy rules, historical patterns, and calendar data to generate suggestions, pre-fills, notifications, and escalation events. Must know its own confidence boundary and defer to humans when data is thin or stakes are high.
Actor · Gallagher
Financial Modelling Analyst / Case Manager
Uses AI-mined source data to prepare cost estimates, feasibility inputs, and scenario comparisons. Must validate the AI output before it becomes a client-facing cost estimate, compliance consideration, payroll output, or approval package.
02 · Starting Point

Starting Point

What is already true when this scenario begins — the world state before AI integration exists.

Starting Point
Assignment data exists, but AI use is not yet structured
The platform may hold assignment records, policy rules, employee profiles, case history, cost-estimate inputs, compliance context, and employee interaction history. The team wants AI to help identify issues, mine the right source data, generate feasibility and cost-estimate inputs, and support consumer interactions, but the boundaries and phasing are not fully defined.
AI Action Mode Reference
Action Type Mode Who Acts Rationale
Field pre-fill on new assignment Assist AI suggests · Human confirms Data errors in assignments have legal consequences — human must verify.
Cost-estimate data mining Assist AI mines · Analyst validates High leverage, but complex financial and compliance assumptions need human review.
Proactive deadline alert Automate AI sends · Human acts Low-risk, high-value — notification delays cost more than false positives.
Draft employee communication Assist AI drafts · Human approves & sends Communications are personal and legally sensitive; AI never sends directly.
Compliance exception detection Escalate AI flags · Compliance Officer resolves High-stakes anomalies require human judgment before any action is taken.
Employee basic question response Escalate AI answers basic · Human handles sensitive Consumer interaction is valuable, but legal, tax, compensation, and immigration answers need guardrails.
Government or payroll submission Assist AI prepares · Human submits Irreversible actions — AI can never execute these unilaterally.
03 · Triggers

Triggers

Five distinct signals that cause the AI layer to act. Each trigger maps to a different mode and makes the platform more proactive without removing human accountability.

Trigger 01 · Automate
Assignment lifecycle milestone is reached
A case moves to a new stage (visa applied, housing confirmed, start date T-30 days, payroll cutoff approaching). The platform detects the transition and the AI layer evaluates which actions are now required, which are overdue, and which stakeholders need to be notified — without any manual prompt.
Trigger 02 · Assist
User begins entering repetitive or inferrable data
A mobility manager starts typing a new assignment form. AI detects a pattern match with 5+ historical assignments for the same route, assignment type, and policy tier. It pre-fills projected fields, suggests the most likely cost estimate, and flags any jurisdiction-specific deviations from the norm.
Trigger 03 · Escalate
Compliance deadline approaches with no action logged
AI monitors outstanding tasks against assignment timelines. When a compliance deadline (visa expiry, tax equalisation due date, home leave entitlement reset) is within 30, 14, or 7 days with no completion event logged, the AI escalates to the responsible owner with full context: what is due, to whom, and what happens if missed.
Trigger 04 · Assist
Cost estimate or feasibility study is requested
AI mining agents gather current source data for a proposed move: rates, policy facts, route assumptions, vendor inputs, prior comparable moves, and feasibility considerations. The analyst or case manager validates the package before any client-facing estimate is released.
Trigger 05 · Escalate
Employee asks a basic or sensitive question
The relocating employee asks about process, status, documents, appointments, compensation, partner work rights, or timing. AI answers approved low-risk questions immediately and escalates sensitive, uncertain, or high-stakes questions to the named human owner with context.
04 · Outcomes

Outcomes

What must be observable when the AI integration is working correctly — the expected behaviours that prove the system is anticipating needs, not just reacting to them.

Expected Outcome Assist
AI pre-fills assignments with >85% field accuracy
When a new assignment is initiated for a known route and policy type, AI auto-populates: projected assignment duration, expected allowances and cost estimates, required compliance steps for the destination jurisdiction, and a draft communication timeline for the employee. Mobility manager reviews and confirms rather than building from scratch.
Expected Outcome Automate
Proactive deadline alerts reach owners before breach
All compliance and operational deadlines have AI-monitored lead times. Alerts fire at T-30, T-14, and T-7 with escalating urgency. Each alert includes what action is needed, who must take it, the downstream consequence if missed, and a direct link to the relevant platform record. Zero deadline misses attributable to lack of notification.
Expected Outcome Assist
AI drafts employee communications; human approves
At each touchpoint milestone (pre-arrival, visa confirmation, housing confirmation, tax briefing due), the AI layer drafts a personalised communication for the relocating employee — pulling in their name, destination, timeline, and relevant entitlements. The mobility manager reviews and sends or edits before dispatch. AI never sends directly.
Expected Outcome Escalate
Edge cases escalate with structured context, not just flags
When AI identifies a case that falls outside policy norms — unusual jurisdiction pair, split payroll, long-term assignment converting to permanent — it escalates to the Compliance Officer with a structured briefing: the anomaly detected, the policy rule being stretched, the risk level (low/medium/high), and recommended resolution paths.
Expected Outcome Assist
Mining agents accelerate cost estimates and feasibility inputs
AI gathers and structures current source data for cost estimates, what-if scenarios, and feasibility studies. The output is available quickly enough to make estimate creation feel near-instant, but it remains a draft until a case manager or analyst validates it.
Expected Outcome Escalate
Employee-facing AI answers basic questions and escalates with context
Employees receive fast answers for approved process, status, document, and appointment questions. When the question touches tax, legal, immigration, compensation, or policy judgment, AI escalates to the human owner with the conversation history and relevant case context.
05 · Guardrails

Guardrails

What must never happen — hard boundaries that protect employees, clients, and Gallagher from AI overreach. Each guardrail names the specific harm it prevents.

Must Not Happen
AI takes irreversible action without human confirmation
AI must never submit a government filing, approve a payroll event, issue a visa application, or trigger a financial transfer autonomously. Any action with legal, regulatory, or financial consequence requires explicit human sign-off. The AI can prepare, draft, and queue — never execute unilaterally.
Must Not Happen
Compliance exceptions are silently suppressed
If AI confidence is low, the exception must be surfaced — not quietly omitted or swallowed. A suppressed exception that later causes a tax or immigration violation is worse than a false positive escalation. The AI must err on the side of visibility, not silence.
Must Not Happen
AI operates on stale or unvalidated data
AI pre-fills, recommendations, and alerts must only fire when the underlying data has been validated and is within its freshness window. A pre-fill based on a superseded policy or an outdated cost rate is more dangerous than no pre-fill. Data provenance and timestamp must be surfaced alongside every AI-generated suggestion.
Must Not Happen
AI-generated cost estimate bypasses validation
No mined estimate, feasibility output, policy interpretation, compliance consideration, or payroll-impacting recommendation can be released as approved without human validation. AI can make the process faster; it cannot replace the case manager or analyst check.
Must Not Happen
Employee-facing AI shares more context than necessary
The consumer interaction layer must not expose internal cost estimates, client-only notes, provider-only data, or sensitive context the employee is not meant to see. AI responses must inherit RBAC and minimum-necessary context rules.
06 · Evidence

Evidence

Observable, measurable signals that prove the AI layer is working. Each evidence card names the measure, target threshold, data source, and what it validates.

Evidence
Pre-fill accuracy measured in UAT across 20+ real routes
Measure: Field-level accuracy rate across a representative set of 20+ historical assignments spanning different route types, policy tiers, and durations.

Target: >85% field-level accuracy overall; tracked separately per field type (cost estimates vs. jurisdictional steps).

Source: UAT comparison of AI pre-fills against known correct values from historical records.

Validates that AI pre-fill is an improvement, not a liability, before production enablement.
Evidence
Deadline alert lead time vs. breach rate tracked monthly
Measure: Time between alert issuance and deadline; whether the deadline was subsequently met or missed.

Target: Zero deadline misses where an alert was issued with >7 days lead time.

Source: Platform event log — alert timestamp, action completion timestamp, deadline date.

Compare breach rate pre- and post-AI monitoring. Report monthly to Implementation Manager and Compliance Officer.
Evidence
Escalation quality scored by Compliance Officer each quarter
Measure: Compliance Officer rates each escalated event 1–5: Was this a true positive? Was the context sufficient to act?

Target: False positive rate <20%; false negative rate <5%.

Source: Structured quarterly review with Compliance Officer using platform escalation log.

Use scoring to tune escalation thresholds. False positives reduce trust; missed positives increase risk.
Evidence
Mined estimate inputs show source, freshness, and validator
Measure: Percentage of AI-mined estimate packages with source links, source dates, confidence, validation owner, and validation timestamp.

Target: 100% before client-facing release.

Source: Cost-estimate mining log and estimate approval workflow.

Validates that mining agents create usable evidence rather than unexplained numbers.
Evidence
Employee AI escalation includes conversation and case context
Measure: Percentage of employee-facing AI escalations that include transcript, question category, case ID, current milestone, and recommended owner.

Target: 95%+ complete escalation packets.

Source: AI chat logs and case escalation records.

Validates that AI reduces repetition instead of forcing the employee to restart with the human team.
07 · Risks

Risks

Three concrete failure modes specific to AI integration — not general platform risks. Each names the harm, the affected actor, and a mitigation path.

Risk
Users over-trust AI outputs; confidence scores hidden
If AI pre-fills and recommendations are presented without confidence indicators or data provenance, mobility managers will treat them as ground truth. A single high-confidence-but-wrong cost estimate or jurisdictional step that slips through unchallenged can result in a compliance breach or budget overrun. Mitigation: surface confidence score and data freshness timestamp on every AI-generated field.
Risk
Training data too thin for rare jurisdiction pairs
The AI model performs well on high-volume routes (UK–US, Germany–Singapore) but may hallucinate or return low-quality pre-fills for low-frequency assignments (Ethiopia–Canada, Peru–Japan). Mitigation: detect low-sample routes and degrade gracefully — show a blank form with a note rather than a low-confidence pre-fill. Log all degraded cases for model improvement.
Risk
Regulatory changes make the model stale without warning
Tax treaties change, visa regimes are updated, jurisdiction-specific thresholds shift. If the AI model is not continuously updated against regulatory feeds, its recommendations silently drift from compliance reality with no user-visible signal. Mitigation: establish a regulatory feed integration and model validation cadence; surface a "last validated" date on jurisdiction-specific suggestions.
Risk
AI scope is priced as one big build instead of phased value
The workshop raised the need for phasing: ideally the full platform exists, but AI-generated estimates, feasibility support, and employee interaction may need to be sequenced. Mitigation: ask engineering for the full breadth, then identify MVP slices by impact, cost, and dependency.
Risk
Employee-facing AI creates unsupported advice
Employees may ask legal, tax, immigration, payroll, compensation, or partner-work questions that sound simple but carry risk. Mitigation: approved answer library, RBAC-scoped context, escalation triggers, and clear human ownership for sensitive topics.
08 · Open Questions

Open Questions

Questions that must be answered before AI features can be locked in and estimated.

Open Question · Product + Legal
Where exactly is the human-in-the-loop boundary?
Is the model copilot (human approves every AI suggestion) or autopilot (AI acts, notifies after)? This needs a decision per action type: pre-fill (copilot), deadline alert (autopilot notification), draft communication (copilot), compliance escalation (copilot). The boundary must be documented, communicated to users, and configurable per client.
Open Question · Engineering + GMS Leadership
Who owns model retraining and feedback loops?
When a mobility manager edits an AI pre-fill, does that edit feed back into the model? Who reviews the feedback queue? Who approves retraining runs? If nobody owns this, the model degrades over time as client-specific patterns diverge from the general training set. A named owner and a retraining cadence must be defined at go-live.
Open Question · Legal + Compliance
How do we satisfy explainability requirements?
Some jurisdictions (EU AI Act) require that automated decisions affecting individuals — including assignment terms, compensation adjustments, or compliance determinations — must be explainable and auditable. Does every AI-generated recommendation need a reasoning trace? What is the minimum explainability obligation for the GMS use case?
Open Question · Product + Engineering
Which AI capabilities belong in the first phase?
The current scenario pack represents the MVP candidate set, but AI-generated cost estimates, feasibility support, policy monitoring, and employee interaction may need phased delivery. Confirm which AI slice should be priced and built first after engineering estimates the full scope.
09 · Decisions & Challenges

Decisions & Challenges

Architectural decisions and workshop assumptions that shape the AI layer.

Decision Required · Product + Engineering
Copilot vs. autopilot mode per action type
Define the mode (copilot / autopilot / hybrid) for each AI action category. Proposed defaults:

Pre-fill → Copilot (human confirms before saving)
Proactive deadline alert → Autopilot (AI sends, human acts)
Draft communication → Copilot (human edits and sends)
Compliance escalation → Copilot + mandatory review before resolution
Government / payroll submission → Copilot (AI prepares, human executes)

This must be locked in before any AI feature ships to avoid inconsistent UX.
Decision Required · Architecture
AI as core platform layer vs. pluggable service
Is the AI layer a first-class platform capability — built into the data model, event bus, and UI — or a pluggable service that clients can enable or disable?

Core embedding: better context, tighter integration, consistent UX — but harder to isolate for regulated clients who may not want AI touching certain data.
Pluggable service: flexibility, client control, isolated by default — but risks fragmented quality and inconsistent experience.

Recommendation: start as a pluggable service with a defined graduation path to core embedding once trust and evidence are established.
Decision · Working Assumption
Include both mining agents and consumer interaction in the AI scope
The workshop explicitly named AI data mining for cost estimates and feasibility studies, and also the consumer interaction point for relocating employees. They may be phased differently, but both should remain in the AI scenario until prioritisation says otherwise.
10 · Score

Scenario Quality Score

The STOA quality engine scores scenarios based on completeness, specificity, and risk coverage. All 9 card types are present with concrete, measurable content.

84
out of 100
Strong — AI scope clarified
5
The Users
1
Starting Point
5
Triggers
6
Outcomes
5
Guardrails
5
Evidence
5
Risks
4
Open Questions
3
Decisions
To reach 90+
Improvement
Define first-phase AI scope
Decide whether mining agents for cost estimates, feasibility study inputs, policy monitoring, or employee interaction should be the first priced AI slice.
Improvement
Specify data-mining source inventory
Name the internal and external sources the mining agents can use, how often each source refreshes, and which sources require human validation before use.
Improvement
Write employee AI answer policy
Define which employee questions AI can answer directly, which need approved content, and which must always escalate to the human case owner.