You are here


There is a myth that shows up in season demos: «shadow mode means delaying the launch». No. In Palma de Mallorca, Ibiza and Menorca, shadow mode is the difference between an agent that proposes and an agent that writes garbage into the CRM on the same day the ferries arrive and cleans run out. Dry-run is not cowardice; it is refusing to let a model invent a check-out or a price while your field team is on the road or on unstable coverage.

If you are planning AI agent app development for multi-island operations, the right order is: the agent sees, proposes and logs; a human reviews; only then it writes. The same applies if you build on mobile apps or custom software in Palma: shadow mode is a product feature, not a «we will look at it later».

Demo, shadow and production

These three stages are not marketing labels. They are different risk levels when a guest in Ibiza writes on WhatsApp at 22:00 and your HQ sits in Palma.

StageWhat the agent doesRisk to guests / opsWhen it fits Balearics seasonWhat you measure
DemoReplies in a sandbox with fake or recorded data.Almost none (no real person receives anything).Winter / pre-season: validate UX and prompts.Does it understand the flow? Does it hallucinate on invented cases?
Shadow (dry-run)Reads real traffic (or a replica) and proposes actions; does not write to customers or CRM.Low if logs are isolated; medium if personal data is filtered poorly.Pre-peak and shoulder: learn ferry, languages and Ibiza/Menorca spikes without touching bookings.% correct proposals, time-to-human review, false positives per island.
ProductionExecutes allowed writes (message, ticket, CRM note) under a permission matrix.High if shadow was short or had no reviewer.Only when shadow is «boring»: rare errors and clear per-island rules.Production errors, escalations, recovery time, internal ops NPS.

If your task map still mixes human and agent without clear borders, review the human, agent or hybrid task map first. Shadow mode does not replace that map: it stress-tests it with real data.

Agent permission matrix

Before you enable writes, fix what the agent may read, what it may propose, and what it must never touch. In Balearic high season, «never» usually includes live dynamic pricing and sensitive personal data without supervision.

System / dataReadWrite (propose → then real)Never
CRMRecords, history, lead/booking status.Notes, tags, handoff card; real writes only after stable shadow.Delete contacts, merge records without a human rule.
WhatsApp / emailInbound threads (with retention and controlled access).Draft replies; real send with a threshold and review.Discount promises, medical data, automated legal threats.
PMS / ticketsAvailability, cleaning states, open tickets.Create draft tickets; change status only in production with rules.Close shifts or mark cleans «done» without sensor/human.
PricingPublished rates, season rules.Suggest ranges in shadow; never apply alone at peak.Change live price without dual OK (ops + revenue).
Personal dataMinimum needed for the task (minimisation).Only fields defined on the card; anonymised shadow logs when possible.Export lists, train on PII without legal basis, share across islands without need.

Good shadow mode also clarifies what the agent must write in the CRM when the human takes over: the handoff card can be trained in dry-run without bothering anyone.

Three typical week-1 failures

These are not yes/no scorecards. They are mistakes we see (and avoid) when shadow mode is treated like a switch, not a process.

  1. Shadow without a human reviewer. The agent fills logs; nobody reads them. HQ in Palma thinks «it is already learning», but nobody flags false positives. Result: in week 3 you enable writes with the same error rate. Assign a reviewer per shift (morning/afternoon) and a daily quota of proposals to review.
  2. The agent proposes without ferry / island context. A «send a tech now» proposal from Menorca ignores the next ferry; in Ibiza it ignores that check-in peaks saturate cleans. Island metadata, lead time and the season calendar must enter the prompt and the rules — not only the WhatsApp text.
  3. Flipping to write mode too early. Two demo days plus one Friday of shadow are not enough. If the painful channel is WhatsApp, map bottlenecks first: the WhatsApp bottleneck map stops the agent from inheriting a broken process. Practical rule: production only when shadow is boring — same exceptions, few surprises.

FAQ

How long should shadow mode run?
It depends on volume. For multi-island ops, a continuous 10–14 day block with real traffic from at least two islands (for example Palma + Ibiza) is usually enough. If Menorca has less volume, extend until you cover weekend spikes.

Does shadow use real customer data?
Yes, or a replica with the same traffic. That is why the permission matrix and PII minimisation matter from day one. Review logs can be anonymised for internal training without losing the ops signal.

Can I shadow on a web app only and later move to mobile?
Yes, if the agent consumes the same events. If the field team lives on mobile, test there too — especially for coverage and offline behaviour in Menorca/Ibiza.

Who decides the move to production?
Not the model. An ops (or product) owner with a signed permission matrix and an agreed error threshold. Marketing and web can wait: they do not replace the OK from whoever pays for the season mistake.

Case: 14 days of multi-island shadow

Picture a field app for cleaning and maintenance crews with HQ in Palma. The proposed agent should: read tickets and messages, suggest priority, draft the handoff note and —later— create the ticket. For 14 days it runs in shadow.

Days 1–3: the agent sees Palma and Ibiza traffic. It proposes priorities; a coordinator in Palma marks hit/miss on a panel. They discover that «urgent» in Ibiza at 21:00 is often keys / access, not an electrical fault. Rules get adjusted.

Days 4–7: Menorca joins. Field connectivity drops at times; the agent starts proposing aggressive WhatsApp retries. The reviewer kills that rule: better a local queue and sync when the network returns. Without shadow, that would have spammed owners.

Days 8–11: the agent settles into a stable proposal pattern (illustrative — not a promised metric). Exceptions cluster: ferry, guest language, apartment without Wi‑Fi. They are written into the matrix.

Days 12–14: shadow is boring. Same exceptions, few new false positives. Only then they enable draft ticket writes — not customer messages. Messages stay as human drafts. That is the controlled jump to production: not «the agent already talks alone».

14-day checklist

  • Define scope: which events the agent reads (CRM, WhatsApp, PMS) and on which islands.
  • Publish the read / write / never matrix and who signs it.
  • Name human reviewers per shift and a daily proposal quota.
  • Log island, ferry/lead time and language metadata on every proposal.
  • Anonymise or minimise PII in training logs.
  • Include at least one relatively busy weekend (Ibiza or Palma) inside the window.
  • Include a poor-connectivity day (Menorca or field) to observe retries.
  • Freeze prompts/rules 48 hours before judging the write go-live.
  • Enable low-impact writes first (CRM note / draft ticket), not the customer channel.
  • Agree the «boring shadow» criterion before celebrating go-live.

Move us to production only when shadow is boring. If you want us to review the dry-run design in your app, send 3 anonymised shadow log examples (proposal + island context) and tell us which island hurts most right now. At Derek Solutions — web design, apps and marketing in Mallorca — reach us at derek@dereksolutions.com · +34 680 283 973.

Contact us now!