Use cases Model EggAI Products Research
Book a demo
EggAI‑v1 · Large Behavior Model

Your customer's digital twin.

A panel that reasons, decides and tells you why — grounded in real transaction histories, and scored against what those shoppers actually did next.

1111101001001111110100011000111011110011101101101010001101010010010000101001110100001101111011111100100110011000001010010101010010101100000000100101010010100000000011110000100010001101101110011001100001111011010101101111110011111010011010011011000110101001111001010010111000001010000111101101011110101000111101111000001110100100100110000011111010110110110111000110011010000010111010110001101000010110101101010011010011010101010100101001110000001110001011101101111101110001000111100111011010100000111011001001100110100011111110111111101001000001110111000011001111111101110110001011101001110110110101111000101001110001011100011101011010011000111111101101010101101101010110000110010100000100100110101011110000111101111100010100100001011100100000101000101011010000111000100101000100001111100000001000010001011011111110011011000111101001001001111101100000011111000110001100011100100110011100100011011001011100000100011000010111011001010010110011101101010100111001111010101000110001000000001101100100010101101010010111000100000101011011100011010011100101010111011001010001110011100001110100110110001000011101100100100010010101011001001101100010110010100010110001101000001110101010110010110001001001001001100100001001111000 11111010010011111101000110001110111100111011011010100011010100100100001010011111101001001111110100011000111011110011101101101010001101010010010000101001110100001101111011111100100110011000001010010101010010101100000000100101010011010000110111101111110010011001100000101001010101001010110000000010010101001010000000001111000010001000110110111001100110000111101101010110111111001111101000000000111100001000100011011011100110011000011110110101011011111100111110100110100110110001101010011110010100101110000010100001111011010111101010001010011010011011000110101001111001010010111000001010000111101101011110101000111101111000001110100100100110000011111010110110110111000110011010000010111011110111100000111010010010011000001111101011011011011100011001101000001011101011000110100001011010110101001101001101010101010010100111000000111000101110101100011010000101101011010100110100110101010101001010011100000011100010111011011111011100010001111001110110101000001110110010011001101000111111101111111101111101110001000111100111011010100000111011001001100110100011111110111111101001000001110111000011001111111101110110001011101001110110110101111000101010100100000111011100001100111111110111011000101110100111011011010111100010100111000101110001110101101001100011111110110101010110110101011000011001010000011100010111000111010110100110001111111011010101011011010101100001100101000001001001101010111100001111011111000101001000010111001000001010001010110100000100100110101011110000111101111100010100100001011100100000101000101011010000111000100101000100001111100000001000010001011011111110011011000111101001001011100010010100010000111110000000100001000101101111111001101100011110100100100111110110000001111100011000110001110010011001110010001101100101110000010001011111011000000111110001100011000111001001100111001000110110010111000001000110000101110110010100101100111011010101001110011110101010001100010000000011011000010111011001010010110011101101010100111001111010101000110001000000001101100100010101101010010111000100000101011011100011010011100101010111011001010010010001010110101001011100010000010101101110001101001110010101011101100101000111001110000111010011011000100001110110010010001001010101100100110110001011011100111000011101001101100010000111011001001000100101010110010011011000101100101000101100011010000011101010101100101100010010010010011001000010011110000010100010110001101000001110101010110010110001001001001001100100001001111000 11111010010011111101000110001110111100111011011010100011010100100100001010011111101001001111110100011000111011110011101101101010001101010010010000101001110100001101111011111100100110011000001010010101010010101100000000100101010011010000110111101111110010011001100000101001010101001010110000000010010101001010000000001111000010001000110110111001100110000111101101010110111111001111101000000000111100001000100011011011100110011000011110110101011011111100111110100110100110110001101010011110010100101110000010100001111011010111101010001010011010011011000110101001111001010010111000001010000111101101011110101000111101111000001110100100100110000011111010110110110111000110011010000010111011110111100000111010010010011000001111101011011011011100011001101000001011101011000110100001011010110101001101001101010101010010100111000000111000101110101100011010000101101011010100110100110101010101001010011100000011100010111011011111011100010001111001110110101000001110110010011001101000111111101111111101111101110001000111100111011010100000111011001001100110100011111110111111101001000001110111000011001111111101110110001011101001110110110101111000101010100100000111011100001100111111110111011000101110100111011011010111100010100111000101110001110101101001100011111110110101010110110101011000011001010000011100010111000111010110100110001111111011010101011011010101100001100101000001001001101010111100001111011111000101001000010111001000001010001010110100000100100110101011110000111101111100010100100001011100100000101000101011010000111000100101000100001111100000001000010001011011111110011011000111101001001011100010010100010000111110000000100001000101101111111001101100011110100100100111110110000001111100011000110001110010011001110010001101100101110000010001011111011000000111110001100011000111001001100111001000110110010111000001000110000101110110010100101100111011010101001110011110101010001100010000000011011000010111011001010010110011101101010100111001111010101000110001000000001101100100010101101010010111000100000101011011100011010011100101010111011001010010010001010110101001011100010000010101101110001101001110010101011101100101000111001110000111010011011000100001110110010010001001010101100100110110001011011100111000011101001101100010000111011001001000100101010110010011011000101100101000101100011010000011101010101100101100010010010010011001000010011110000010100010110001101000001110101010110010110001001001001001100100001001111000 MILLIONS OF REAL CUSTOMERS EGGAI-V1 MILLIONS OF DIGITAL TWINS
Case study 01 Offer optimization

Pay less for the same yes.

DMBGN Lazada voucher benchmark · SIGKDD 2021 · BSD‑2 licensed

Most discounts are bigger than they had to be. We ask each customer’s twin how far the offer can come down before they walk away — then price to that, one customer at a time.

HOW MUCH OF THE DISCOUNT BUDGET WE TAKE BACK 6% PROVEN 15% TYPICAL 0% 5% 10% 15% 20% FLOOR IS WHAT REAL TRANSACTIONS ALREADY BACK — THE MODEL ITSELF REACHES FURTHER, AND WE PROVE THAT ON YOUR OWN CAMPAIGN
6–15%
off the discount budget
0%
lost in sales
939
offers re-priced
Measured on a real campaign of 939 offers. Sales stay where they were — the customer still says yes, only the discount comes down. The floor is what real transactions already back; we prove the rest on your own campaign.
Case study 02 Lead generation

Stop calling the people who say no.

Same benchmark, read the other way round

You will never call the whole list — there are always more leads than hours. So the only question that matters is which slice your team spends its day on.

20 PEOPLE ON YOUR LIST · 3 OF THEM WOULD SAY YES YOU HAVE TIME FOR 4 THE ORDER YOU GOT THEM IN YOUR FOUR CALLS REACH NOBODY RANKED BY THE MODEL THE SAME THREE BUYERS, NOW IN REACH GREEN = SOMEONE WHO WOULD HAVE SAID YES · THE ORDER IS ILLUSTRATIVE
15
in 100 on the list ever buy
76
in 100 put in the right order
12,480
real offers tested on
Tested on 12,480 real offers where we already knew who bought. Bring your own list and we will show you which end to start from.
Case study 03 Survey response

Answer like your customers do.

Health & Beauty panel · 27 questions · N = 1,013

An outside agency had already asked 1,013 people in Thailand whether they stick to one brand. We put the same question to a digital twin panel.

“DO YOU STICK TO ONE BRAND?” REAL PEOPLE 38.0% DIGITAL TWIN PANEL 41.3% UNTRAINED ENGINE 10.0% FRONTIER AI 8.7% OUT OF 10 SHOPPERS, HOW MANY STICK TO ONE BRAND BODY WASH CATEGORY · 188 REAL BUYERS
3
points off the real answer
64%
match across all 27 questions
52%
for a general-purpose AI
One of 27 questions from the study. Across all 27 the panel lands 64% of its answers where the real people did, against 52% for a general-purpose AI.

Try it on your own category.

Forty-five minutes on a category you know well, with the panel answering live and the reasoning visible for every answer.

Why now

Agents learned to reason. They never learned to shop.

A model can only get better than the humans it copies if something scores its answers. Retail already keeps that score.

01
Agents need a world
Act, see what happened, try again. Code and maths got their sandbox years ago; marketing never had one.
02
Today we only copy humans
Imitation learning and preference tuning both stop where human judgement stops. So does every synthetic persona built on them.
03
Verifiable reward lifts the ceiling
Baskets, redemptions and repeat visits are ground truth. Score against them and the model can pass the people it learned from.
Use cases

Two ways to point one model at a list.

Both ask the same question of one customer at a time — would you accept this? — and both are graded against offers real people did or did not take.

Use case 01
Offer optimization
Every offer you send has a price you did not need to pay. Propose the stingiest terms first, read which constraint the twin objected to, move only that one, and stop at the cheapest offer this customer still accepts.
Puts the likely buyer ahead of the unlikely one
76 times in 100, across 12,480 real offers
The measurement →
Use case 02
Lead generation
The same engine, pointed at a list instead of an offer. Score every name for predicted acceptance, rank, and cut the tail — so an SDR works the top of a list instead of all of it. The model is deliberately hard to please.
Out of 12,480 offers it approved 120
the rest came back "not worth the call"
The measurement →
Products

Three SDKs, one behaviour model.

The case studies are what the model does. These are how you call it.

Offer Optimization SDKOffer optimization
Score a proposed offer, or search down to the cheapest terms a customer still accepts — with the constraint they objected to.
Lead Generation SDKLead generation
Send a list, get it back ordered by predicted acceptance, with a reason against each name and an explicit not-worth-calling verdict.
Survey Response SDKSurvey response
Put a questionnaire to a simulated panel and get a distribution of answers back, with the reasoning attached.
The model

EggAI‑v1

EggAI‑v1 answers as one named shopper rather than as a segment. It is built on the real till history of a major SEA grocery retailer — every trip in the order it happened — and on a training mix calibrated to how people in this region actually answer. Every reply is scored against what those shoppers really did next, so it can be marked right or wrong.

See how it was tested See the two use cases
Parameters
27B
trained to answer as a person, not prompted to
Trained on
4.07M transactions
1.02M real baskets, read in the order they happened
Price accuracy
82.0%
on what they said they would pay · frontier model 67.1%
Offer ranking
76 in 100
redeemer above non-redeemer · 12,480 held-out offers
Inside one answer

Hand it a profile. It thinks as that person, then decides as them.

The shopper it was given
Gets to the shop about 1.5× a month
Median basket ฿241
Uses a coupon on 14% of trips
Some discount on 86% of trips
Buys confectionery now and then
The offer put to them
15% off confectionery
minimum spend ฿200
What it thought, in the first person
“I use coupons quite a lot, actually — I only get to the shop about one and a half times a month, so I have to find the way that saves me the most. And I buy sweets fairly regularly anyway. So I’ve decided to just go ahead and use this 15% coupon.”
Reasoned in Thai and translated here — it answers in the shopper’s own language, in the register they would actually speak.
It decided
Accept
What they really did
Accepted
Nothing here was retrieved from a lookup table. The model was shown this shopper’s own history, reasoned about the offer the way they would, and landed where they landed — and the reasoning is what tells you why, so you can argue with it before you spend a campaign on it.
Training architecture

Three phases, and a reward that comes from what shoppers really did.

Phase 01 · Pre-train
Become a regional shopper
A general model is taught to be conditioned into one particular SEA shopper. A customer’s life is read as a sequence of occasion‑then‑basket turns — what the trip was for, what came home — with the reasoning written in that customer’s own voice, in their own language.
Phase 02 · Fine-tune
Every task is a survey
Each task is framed as one question put to one person, which is what keeps a score and a stated reason attached to the same answer. Where a teacher model writes the reasoning, the answer is pinned to the known ground truth first and only the why is generated — so a persuasive rationale can never talk the model into the wrong label.
Phase 03 · Reinforce
Match the spread, not the mode
The reward is for matching the distribution of real answers, never for sharpening onto a single confident one. A simulated panel that agrees with itself more than real people do is wrong about people in a way accuracy alone will not catch, and this is the phase that holds it apart.
The evidence

Three fielded studies stand behind these numbers.

Each one is a real study with a real topline the model was never trained on — the questionnaire, what it was shown, what it was kept from seeing, and how every figure was worked out.

Read the three studies
Products

One model, three SDKs.

EggAI ships as libraries you call from your own systems. Each one points the same behaviour model at a different decision, and each has a published study behind it.

01 Offer optimization

Offer Optimization SDK

What is the least this customer would accept?

A weekend promo is ready at 15% off produce, and finance wants the discount smaller without losing redemptions. Ask each customer’s twin what the cheapest version they would still take is — 8% for one, a lower minimum spend for another, nothing at all for a third — and send each of them that, instead of one blanket 15%.

The research behind it →
Call it like this
from eggai import OfferOptimizer optimizer = OfferOptimizer(api_key) terms = optimizer.cheapest_accepted( customer = "c_8813", profile = "shops fortnightly, small baskets, " "coupon-led, buys confectionery", offer = "Fresh Picks Weekend — 15% off " "produce, min spend 500", guidelines = """ discount caps at 12%, min spend above 300. the produce section is fixed. the tagline may be reworded, but never promise free delivery. """, ) terms.text # "8% off produce, min spend 349" terms.tagline # "Weekend Picks" terms.changed # ["discount", "min_spend", "tagline"] terms.objection # "threshold above my usual basket"
Removes 6–15% of the discount committed to a re-priceable campaign, holding acceptance fixed.
02 Lead generation

Lead Generation SDK

Which of these people is worth the call?

Four thousand names, six people on the phones, two days before the promo closes. Rank the list so the team opens with the calls most likely to land and stops where the next call is not worth the minute — and, if you want it, get the version of the offer each name is likeliest to say yes to.

The research behind it →
Call it like this
from eggai import LeadGenerator customers = { "c_8813": "fortnightly, coupon-led, small baskets", "c_2291": "weekly family shop, rarely uses vouchers", } leads = LeadGenerator(api_key) ranked = leads.rank( customers = customers, offer = "Fresh Picks Weekend — 15% off produce", guidelines = "hold the 15%, the aisle may change", tailor_offer = True, # False = score and rank only ) ranked[0].score # 0.71 ranked[0].reason # "buys dairy most weeks, produce rarely" ranked[0].offer # "Fresh Picks Weekend — 15% off dairy" ranked.above(0.5) # the slice worth calling
Orders a buyer above a non-buyer 76 times in 100 across 12,480 held-out offers.
03 Survey response

Survey Response SDK

What would a panel of your customers say?

You need to know whether shoppers would move to a new own-brand line, and fielding a real panel is three weeks and a budget line. Put the questionnaire to 150 simulated customers drawn from real purchase histories, read the spread of answers today, and field the real panel on the questions that still look uncertain.

The research behind it →
Call it like this
from eggai import SurveySimulator survey = SurveySimulator(api_key) result = survey.ask( # or customers = {id: profile} for your own list panel = "shoppers in Thailand, 18-55, who buy " "fresh produce most weeks", n = 150, question = "Would you switch to our own-brand line?", options = ["Definitely", "Maybe", "No"], guidelines = "let the unsure stay unsure", ) result.distribution # {"Definitely": 0.21, "Maybe": 0.44} result.transcripts # the reasoning behind each answer result.crosstab("age") # the same spread, split by segment
Lands 64% of its answers where 1,013 real respondents landed theirs, against 52% for a frontier AI.

What you need to start

Two of the three run on your transaction history. The third needs nothing at all.

SDK
You provide
Runs as
Offer Optimization
Transactions and your promotion calendar
Batch or per-call
Lead Generation
Transactions, and the list you want ranked
Batch or per-call
Survey Response
Nothing — ask and it answers
Per-call
Research

Three studies, written up in full.

Each SDK rests on one of these. Every paper gives the data, the split, the scoring and the result — including where a purpose-built system beats us.

Study 01

Pricing each offer to the cheapest terms a customer still accepts

A per-customer search over discount depth and minimum spend, applied to a completed voucher campaign.

Discount removed6–15%
Offers re-priced939
Sales given up0%
Read the paper →
Study 02

Ordering a contact list by predicted acceptance

Ranking quality on a held-out voucher benchmark, and what it implies for a team that cannot call everyone.

Pair ordered correctly76 in 100
Held-out offers12,480
Base acceptance15.3%
Read the paper →
Study 03

Replicating a fielded consumer panel with simulated respondents

A 27-question health-and-beauty study, re-run against a simulated panel and marked on distribution match.

Answer match64%
Frontier AI52%
Real respondents1,013
Read the paper →
The position

The ceiling comes off when the reward is real.

Why a behaviour model beats a bigger general one at this, in three moves.

01
Agents need a world to act in
Every capability jump of the last three years came from a sandbox with a scorer in it — a compiler, a unit test, a game. Marketing never had one. Plans are argued, shipped, then explained after the fact.
02
Copying humans has a ceiling
Imitation learning reaches the average of its demonstrations; preference tuning reaches the taste of its raters. That is why prompted personas feel plausible and predict badly — they summarise how people say they shop.
03
Retail already keeps the score
Baskets, redemptions and repeat visits are recorded, dated and unarguable. Train against them and reward stops being a matter of opinion: the model is graded on what a real person did next.
What we publish
The method, in full: how each evaluation is built, what the model is shown, what it is deliberately kept from seeing, and how every number was worked out. Under an NDA we will re-run any figure on this site in front of you.
What we don’t
It learned from a client’s real till data, protected under Thai privacy law, and the trained model carries that with it. Neither leaves our own servers, and we would rather say so than look more open than we are.
Brand

The mark is a signal in a shell.

An egg drawn twice — amber behind, emerald in front — with one trace running through it to a node. The offset is the echo of a customer; the trace is what we read.

Primary mark
Horizontal lockup
EggAI
On light
EggAI
Stacked
EggAI
Mono ink
Mono ground
App icon

Palette

Two brand colours, read as a spectrum. Everything else is ground and ink.

Amber
#FFC700
Emerald
#00C88C
Ink
#F3F8F3
Ground
#040504
#0A0C0A surface
#1B2119 line
#8F998B muted
Spectrum
Typography
Display 500
Headlines · −4.5% tracking · 46–86px
Title 600
Wordmark, card titles, figures · −3% tracking
Body 400 at 16–20px, set to 1.6 line height and never wider than 62 characters.
Paragraphs · #A3ADA0 on ground
Label 500 · uppercase · 16% tracking
Rules
DoKeep clear space of one egg-width on every side. Never below 24px tall.
DoLet the spectrum run amber to emerald, left to right or top to bottom.
DoUse one glow per screen, behind the first thing the eye lands on.
NoDon't fill the mark, add a stroke weight of its own, or rotate it.
NoDon't swap the two colours — emerald is always the front egg.
NoDon't set the spectrum gradient behind body text.
Contact

Bring a brief. Leave with an answer.

Forty-five minutes on your category, with the panel answering live. No slides unless you ask for them.

Name
Work email
Company
Category
What would you ask the panel?
We reply within a working day.
Direct
Saleshello@eggai.com
Researchresearch@eggai.com
Presspress@eggai.com
Where we are
Bangkok
EggAI, an Amity company
Sathorn, Bangkok 10120
Thailand
No weights to download and no leaderboard to point at. What we can do is run your own category through the panel while you watch, and show you the numbers where we lose as well as the ones where we win.
Use cases

Two ways to point one model at a list.

Both read one customer’s own order history and answer a single question in first person — would you accept this? One walks the terms down until the answer turns to yes. The other keeps the names where the answer is already yes and drops the rest.

Use case 01

The cheapest offer they would still accept.

A discount that gets redeemed tells you nothing about the discount you could have got away with. The model answers the counterfactual one customer at a time: at these terms, would this person still redeem?

So the search runs downward. Start at the smallest discount and the hardest minimum spend, read the constraint the twin actually objected to, move only that lever, and stop the moment the answer turns.

How the search runs
01Build the twin from that customer’s own ordered history.
02Propose the stingiest plausible terms — smallest discount, hardest minimum.
03Ask in first person. Get a likelihood, a yes or no, and the reason.
04Move only the lever the reason named. Re-score. Up to eight rounds.
05Stop at the first acceptance — or return not targetable at this cost.
Unedited model output
One shopper, one public-benchmark offer. Same 500-unit discount both times; only the minimum spend changed.
500 off · min 3,999
NO
500 off · min 299
YES
“I’m likely to redeem this voucher because my recent activity shows I’m already in a voucher-usage mindset, and a 500-off voucher with a 299 minimum spend fits my usual high-value, voucher-driven shopping pattern. Since I…”
The reasoning is the model’s own, in first person. It is a prediction of this shopper’s behaviour, not a record of it. The threshold moved the decision; the discount did not.
Use case 02

Cut the list before anyone calls it.

The same engine, pointed at a list instead of at an offer. Score every name for predicted acceptance, rank them, and cut the tail — so an outbound team works the top of a list instead of all of it.

The useful part is how hard it is to please. Asked about 12,480 real offers it said yes to 120 — fewer than 1 in 100. Everything else came back as not worth the contact. A short list someone can defend cutting is worth more to a rep than a long one nobody finishes.

What the scoreboard says
76 in 100
Hand it one person who used their voucher and one who ignored theirs. It puts the two in the right order 76 times in 100, across 12,480 real offers where the answer was already known. Pure chance is 50. Shown the shopper’s live visit as well, it gets to 83.
120 of 12,480
How much of that list it was willing to keep — fewer than 1 name in 100. Everything else came back as not worth the call.
15 in 100
How many it expected to accept across the whole list. About 15 in 100 really did — so the overall rate is right, even where its confidence in one person is not.
What it needs
Purchase history for each name on the list. The model scores a person from what they have bought, so it works on a customer base, a loyalty file or an installed base — not on a name with no history behind it.
Cold names
On 414 people it knew nothing at all about, it got 71.7% of buying decisions right. A general-purpose AI got 70.6%. Treat that as a tie: with no history to read, there is nothing for it to be better at.
Order, don’t forecast
When it feels sure about someone it is over-optimistic, by roughly two to three times. So trust the order it puts your list in. Do not read the percentage beside a name as a forecast until it has been checked against your own campaign.
How we tried to prove ourselves wrong

Both use cases rest on the model preferring some offers to others, so we tried to trip it up. We slipped a meaningless sentence into the question — a line about Thailand having 77 provinces — and the score moved a little. Then we personalised the actual offer, and the score moved about the same amount, while the number of people who said yes did not budge: 120 before, 120 after.

So the honest reading is that this is a sorter, not a persuader. It orders a list well enough to cut it, and it finds the point where someone stops saying yes. It has not shown that changing an offer causes a person to accept. Proving that means holding back a random group and deliberately not sending to them, and we have not done it.

Everything on this page is a prediction, marked against what people really did under offers somebody else had already chosen. Where a number would need a test we have not run, it is not on the page.

Bring a list. We will tell you what to cut.

Forty-five minutes on your own category, scoring your own names, with the reasoning visible for every one.

Book a demo
Study 01

Pricing each offer to the cheapest terms a customer still accepts

A per-customer search over discount depth and minimum spend, applied to a completed voucher campaign.

EggAI Research · behaviour-model line · public benchmark data
Abstract

Discounts are set for segments, so most customers receive more than they needed to convert. We ask a behaviour model, customer by customer, how far an offer can be reduced before that customer stops accepting it, and re-price to the point just above the refusal. Applied to 939 re-priceable offers from a completed campaign, the procedure removes between 6% and 15% of the discount committed to them while holding predicted acceptance fixed.

1Background

An offer that is redeemed tells you the terms were sufficient. It does not tell you whether they were necessary. The counterfactual — would this customer have accepted less? — is unobserved for every redemption in a campaign, because only one set of terms was ever sent.

A behaviour model trained on real purchase histories can be asked that counterfactual directly, one customer at a time, at terms that were never offered.

2Data

Source
Detail
Benchmark
DMBGN Lazada voucher redemption, SIGKDD 2021, BSD-2 licensed
Held-out set
12,480 offers, scored in full
Redemption rate
1,909 of 12,480 — 15.3%
Re-priceable subset
939 offers for which a cheaper valid term existed
Contamination
Zero rows from this dataset appear in training, audited
Terms are parsed from the offers as issued. The publisher desensitises the currency, so amounts are the benchmark’s own units.

3Method

For each customer the search begins at the least generous plausible offer — smallest discount, highest minimum spend — and asks the twin, in the first person, whether it would redeem. The reply carries a likelihood, a decision, and the constraint it objected to.

Only the lever named in that objection is moved, and the offer is re-scored. The search runs up to eight rounds and stops at the first acceptance, which is that customer’s cheapest accepted terms. Where nothing inside the budget is accepted, the procedure returns no offer rather than manufacturing one.

Search, one customer
start smallest discount, highest minimum spend ask “would you redeem this?” → likelihood, yes/no, objection move only the lever the objection named repeat up to 8 rounds stop first acceptance = cheapest accepted terms

4Results

Share of the discount committed to the 939 re-priceable offers that is removed, holding predicted acceptance fixed.
Basis
Discount removed
Corroborated by real transactions
6%
Where the model is confident
42%
Every offer at its floor
58%

The published range is 6–15%: the floor is the band corroborated by what comparable customers actually did, and the upper figure is set well inside what the model itself reaches. Sales are unchanged in every band, since the procedure holds acceptance and moves only the discount.

Measured against the campaign’s full discount bill rather than the re-priceable subset, the same bands are 1.7% and 16%. The subset carries 27% of the bill.

5Discussion

The binding constraint is reach, not pricing. The 939 re-priceable offers hold 27% of the campaign’s discount; the rest sits on offers the model either did not identify as redemptions or could not price lower. Improving recall on high-value offers moves the ceiling further than improving the search.

Every figure ranks or matches behaviour recorded under the campaign that was actually run. Establishing that cheaper terms cause the same purchase requires a randomised holdout cell, which is the natural next step for a live deployment.

All results Next study →
Study 02

Ordering a contact list by predicted acceptance

Ranking quality on a held-out voucher benchmark, and what it implies for a team that cannot call everyone.

EggAI Research · behaviour-model line · public benchmark data
Abstract

Outbound teams hold more leads than they have capacity to contact, so the operative decision is which subset to work. We evaluate a behaviour model as a ranker on 12,480 held-out voucher offers with known outcomes. Presented with one person who accepted and one who did not, the model orders the pair correctly 76 times in 100 against a chance rate of 50, and at its own acceptance threshold it retains under 1% of the list.

1Background

A list is worked from the top. If its order is arbitrary, the acceptance rate of the first hour equals the acceptance rate of the list, and on this benchmark that is 15 in 100 — the remaining 85 contacts reach someone who was never going to accept.

Ranking is therefore the whole product: not whether a contact will accept, but whether they will accept sooner than the next one.

2Data

Source
Detail
Benchmark
DMBGN Lazada voucher redemption, SIGKDD 2021, BSD-2 licensed
Split
Publisher’s own: 49,588 train, 12,480 held out
Scored
The full held-out set, not a sample
Positives
1,909 redemptions — a 15.3% base rate
Contamination
Zero rows from this dataset appear in training, audited

The model was trained on a different company, in a different country, answering a different question. Each row carries a shopper’s profile and their activity before the voucher was collected; activity after collection is excluded, since it is unavailable at decision time.

3Method

Each row is scored by reading the probability the model assigns to acceptance directly, which is the readout the published benchmark uses. Ranking quality is the probability that a randomly drawn accepter is placed above a randomly drawn non-accepter.

One held-out row
System You are an online shopper in Southeast Asia. Demographics: age level 5.0/8, female, purchase level 8.0/10. History: 69 orders, 9 used a voucher (13% voucher-usage rate), average order 8.7 (range 0.6–60.2). User Your activity around this voucher: — BEFORE collecting it: added 4 items to cart, placed 3 orders. — You browsed 6 distinct product categories in this window. Campaign “C2”: voucher = 500 off, minimum spend 4,999. Will you redeem it? Answer exactly YES or NO. Gold NO (the recorded outcome)

A second readout — asking the model to reason in the first person, then judging the written answer — was run on the same 12,480 rows. It is the readout used wherever a closed frontier model appears in a comparison, since such a model will not expose a probability.

4Results

System
Pair ordered correctly
EggAI-v1
76 in 100
A system built only for this task
79 in 100
An ordinary statistical model
79 in 100
Chance
50 in 100
Direct-probability readout, full 12,480-row held-out set.

On the judged-answer readout, where a frontier model can be included, the same comparison is 69 in 100 for the behaviour model against 64 for a frontier model and 60 for the same engine untrained — with no examples from the target platform at any point.

The model is also markedly selective. At its own acceptance threshold it returns yes on 120 of the 12,480 offers, under 1 in 100, and its mean predicted acceptance across the set is 0.155 against a true rate of 0.153.

5Discussion

A purpose-built system and a conventional statistical model both reach 79 on this benchmark, three points above the behaviour model, which suggests the remaining gap is a property of the available features rather than of the approach.

The split is disjoint by session rather than by customer, so the result characterises ordering among contacts a business already has history for. Scoring a contact with no history is a separate problem: on a cross-market test of 414 consumers with no identity overlap with training, the model reached 71.7% on purchase decisions against a frontier model’s 70.6%.

All results Next study →
Study 03

Replicating a fielded consumer panel with simulated respondents

A 27-question health-and-beauty study, re-run against a simulated panel and marked on distribution match.

EggAI Research · panel-model line · client study, reported in aggregate
Abstract

Survey research is slow and expensive to field, and a simulated panel is only useful if it reproduces the spread of a real one rather than a single plausible answer. We took a health-and-beauty study already fielded to 1,013 people and put the same 27 questions to a 150-respondent simulated panel. Across all 27 questions the simulated panel places 64% of its answer mass where the real respondents placed theirs, against 52% for a general-purpose frontier model and 47% for the same engine without behavioural training.

1Background

A model asked a survey question will produce a plausible answer. A panel produces a distribution — a spread of answers across a population, with minorities and fence-sitters in proportion. Reproducing the first is easy and commercially useless; the second is the thing a research buyer is paying for.

The evaluation therefore marks distributions against distributions, question by question, rather than scoring any individual answer as right or wrong.

2Data

Source
Detail
Real study
A health-and-beauty consumer study fielded by an outside research agency
Respondents
1,013 people in Thailand, 18–55
Instrument
27 single-answer questions: spending trend, out-of-stock reaction, brand loyalty across 11 sub-categories, format and pack size, channel split
Simulated panel
150 respondents, matched to the same 18–55 screener
Contamination
No model in the comparison had seen this study, its questionnaire, or its results
The client is not named, and no brand, pack size, channel or price from their data appears here. Results are reported in aggregate.

3Method

Every model receives the same questionnaire in the same format and is marked by the same analyser. For each question, the simulated panel’s answer distribution is compared with the real topline; the reported figure is the share of answer mass that lands in the same place, averaged over the 27 questions.

Elicitation is held constant across models, since the mode in which a panel is asked to answer materially changes the spread it produces.

4Results

Panel
Answer match
Same top answer
The 1,013 real respondents
EggAI-v1
64%
63 in 100
A frontier AI
52%
48 in 100
The same engine, untrained
47%
15 in 100
Mean across all 27 questions. The behavioural panel is closest on 11 of the 27; the frontier model on 6.

On the brand-loyalty question in the body-wash category, where 188 real buyers answered, the real study found 38.0% keep to a single brand. The simulated panel returned 41.3%, the untrained engine 10.0% and a frontier model 8.7%.

The frontier model’s characteristic failure is visible on the spending-trend question, where it placed 90.7% of its respondents on one middle option. Ten percent of real respondents chose it.

5Discussion

The gap between the untrained engine at 47% and the trained panel at 64% is attributable to behavioural training rather than to model scale, since both share the same base. The frontier model sits between them at 52% while collapsing on the questions where a population is most dispersed.

Distribution match is an aggregate property. Respondent-to-respondent variation within the simulated panel remains narrower than in a fielded one, which places the instrument at questionnaire piloting, concept screening and segment sizing rather than at replacing fieldwork.

All results Next study →
EggAI
An Amity company. Building the behaviour simulator that agents needed all along.
Product
EggAI‑v1 Offer Optimization SDK Lead Generation SDK Survey Response SDK
Research
Use cases Research
Company
Contact Brand kit
© 2026 EggAI · an Amity company Every figure on this site is measured. Where we lose, we say so.