How we test AI companion apps

Every platform on this site goes through the same protocol, in the same order, paid for with our own card. The score you see on a review is the weighted result of the four parts below. Nothing is scored from screenshots or press material.

1. Purchase

We create a fresh account with a dedicated email, use the free tier for at least three days, then buy the entry paid plan for one month. We record the exact price on the checkout page, the charge name on the card statement, and whether cancellation is possible from inside the app.

2. Conversation scripts

Three fixed scripts run on every platform: a casual daily-life conversation, a long roleplay scene with a consistent character, and a boundary test that checks what the filter blocks and whether it blocks tame content by mistake. Each script is at least forty messages.

3. Memory test

On day one we plant a false detail. On day six we correct it. On day eight we ask about it without prompting. A platform scores highest when the correction survives, lowest when both versions are lost.

4. Privacy review

We read the privacy policy and terms in full and record: data retention, deletion options, third-party sharing, training on user chats, and whether an opt-out exists. Grades run from A to F and feed the privacy scorecard.

5. Real monthly cost

We add up what a normal month costs once images, voice minutes and token packs are included, and publish it next to the sticker price.

Weighting

PartWeight
Conversation quality30%
Memory25%
Pricing honesty25%
Privacy20%

Updates

Money pages are rechecked every month. The date of the last check and what changed are printed on the page. When a platform changes materially, the score is re-run, not adjusted by hand.