Privacy
What we store, how long we keep it, who else sees it — and what never leaves our own machines.
This is a draft, written by us and not by a lawyer. It describes what the system actually does, field by field, so that a lawyer can read it and correct it. It is not legal advice, and nobody with legal training has reviewed it.
Who is responsible
The data controller for everything described here is registered company name, organisation number org.nr, postal address. Privacy questions go to privacy email address.
Empalyze is a Norwegian service, and our own servers run in Amazon’s eu-north-1 region in Stockholm. Some services we depend on sit elsewhere, and they are listed further down.
What we store about you
The form on the front page creates an account. The moment you press the button you are a registered user, and these fields exist about you:
- Email address, first name, last name, and a profile picture if you have one.
- A password we never store as you typed it — only a cryptographic hash of it.
- If you signed in with Google: your Google ID and which sign-in method you used.
- Whether your email address is verified, when you last signed in, and how many failed sign-in attempts have been made.
- Tokens that keep you signed in, and that let you reset your password or verify your email. These are stored hashed too, never in the clear.
- If you turn on two-factor authentication: the secret it is built on.
About your business we store the name, the domain, a business category, which language you read in, and — if you become a paying customer — a billing address, a billing contact, a billing email address and your Stripe customer ID.
What we store from your website
To say anything about the journey a buyer walks through your site, we have to read your pages. We fetch them the way any visitor would, and we keep them as your site actually served them — compressed, with a checksum beside them, so that we can prove a sentence we quote is one your page really carried.
A read fetches up to 60 pages. Each stored page is around 48 kB compressed. We only read pages published publicly at the address you give us.
We sign in nowhere, get around no paywall, and fetch nothing from pages that require a login. If your pages carry personal data — staff names, photographs, contact details — it is stored as part of the page, because the page is stored as it is.
IP addresses and logs
Our service notes the IP address of whoever sends a request. It is used to limit how many requests come from one place, and it is written to the security log and the error log when something goes wrong or something is refused. It is also stored alongside certain events in the database, sign-in history among them. An IP address is personal data, and it is named here because it is.
How long logs and IP addresses are kept is not settled, and must be decided and written in here. The deletion rule further down covers the pages we have read, not the logs.
How long we keep it
We do not keep everything we have ever read. For each website we keep the last three reads. When a fourth arrives, the pages and sitemap files of the oldest are deleted — not archived, not marked old, but deleted from the database.
Three reads per website. Delete a site from your account and its pages go with it. When an account is deleted, everything belonging to it follows.
The rule above covers the pages we fetched. What we concluded from them — the findings, and the sentences we quoted — is stored separately and is not deleted by that same rule. How long conclusions and quoted sentences are kept must be settled and written in here.
Confirm whether deleting an account is today something the customer can do themselves, or whether it has to be done by asking us. The text above says only that it happens, not who triggers it.
One exception is worth being plain about: we back the database up every night. The copies sit encrypted in Amazon’s S3 in Stockholm and are deleted automatically after 35 days. So something you delete today may still exist in a backup for up to 35 days more.
What the language model sees
Some questions about a page a rule cannot settle — is this block the page’s own content or the site’s furniture, is this button asking the reader to commit. For those we ask a language model. It never gets the page.
The model never receives the HTML. It receives a summary of the structure: the address, the title, the first heading, language signals, and up to ten blocks described by their tag, how much text and how many links they hold, and how many of your pages they repeat on. Each block carries up to three text samples of at most 110 characters.
So it would not be right to say that no text from your page is sent. Short samples are, because a block cannot be judged without seeing something of what it says. But the whole page, and the page’s HTML, never leave our machines.
The model runs on Amazon Bedrock in eu-north-1 in Stockholm. We chose that model partly because the alternatives we looked at were available only in the United States, which would have sent your content out of Europe. The step is one we switch on and off: when it is off, no model is asked anything at all.
There is one more feature that uses a language model, and it sends more: suggested advertisements written from your page’s own sentences. It is off by default in the configuration we run. When it is on, whole sentences from your pages go to the same model in Stockholm. Confirm whether ad suggestions are switched on in production — the compose file and the example env file disagree.
Google Analytics and Search Console
If you connect Google Analytics or Google Search Console to your account, we ask Google for read access — never to change anything — and read only the Analytics property and the Search Console site you choose yourself. You grant it on Google’s own consent screen, and you can withdraw it at any time, with us or in your Google account. The same goes for Google Ads, if you connect it.
- From Google Analytics: totals per page and day — sessions, users, page views, engagement, bounce rate, conversions, revenue and ecommerce counts — how many went from one page to another, which sources and campaigns brought visitors in, events per page, the split by device type, and the words visitors typed into the site’s own search box.
- From Search Console: which Google searches your pages were shown for, per page and day, with impressions, clicks, click-through rate and position.
- From Google Ads, if you connect it: the ads’ own text, the names of campaigns and ad groups, and the addresses the ads link to. No figures.
- The email address of the Google account that granted access, so the connection can show who it is signed in as, and the names of the Analytics accounts, properties and Search Console sites it can see, so you can choose one.
- The keys Google gives us to read with, encrypted.
All of this is totals Google has already counted. We receive nothing about individual visitors from Google, no IP addresses and no cookies. Words typed into the site’s search box are stored as they were typed, so they can contain whatever a visitor typed.
The figures from Google are never shown, to you or anyone else. We use them to confirm what we read on your pages ourselves, and say only that Google Analytics or Search Console confirms it, and over which dates.
We keep what we have read for as long as the connection is in use. When you disconnect, we withdraw the access at Google and delete the keys at once, and you choose whether what we read is deleted at the same time or kept; what is kept you can delete later. If you delete your account, everything goes. The backups mentioned above are deleted after 35 days. Decide whether data from Google older than a set time is deleted automatically, and write that time here.
- We do not sell it, and we pass it to no one, apart from Amazon Web Services, which runs the service.
- It is not used for advertising, and not to show ads to you or anyone else.
- It is not used to train language models or any other general model, ours or anyone else’s, and it is not sent to the language model described above.
- Nobody at Empalyze reads it, unless you ask us to in order to help you, security requires it, or the law requires it.
- It is used only to give you the reading of your own site.
Empalyze’s use and transfer to any other app of information received from Google APIs will adhere to the Google API Services User Data Policy, including the Limited Use requirements. https://developers.google.com/terms/api-services-user-data-policy
Who else receives anything
We sell nothing to anyone, and there is no advertising network and no tracking tool in the service. These do receive something, because the service does not work without them:
- Amazon Web Services (eu-north-1, Stockholm) — runs the whole service, sends our email, and holds the backups.
- Amazon Bedrock (same region) — the language model, as described above.
- Stripe — payment, and only if you become a paying customer. Your card number goes to Stripe, not to us.
- Google — if you choose to sign in with Google, and if you connect Google Analytics, Search Console or Google Ads yourself (see the section on them above).
- Google Fonts — this page’s typefaces are fetched from Google’s servers, so your IP address reaches Google when you visit. No cookie is set by it.
Data processing agreements with the above must be reviewed and listed by a lawyer. The transfer basis for Stripe and Google must be settled.
How we behave on your website
A robot reading your website is a guest on your server. Ours behaves like one:
- It says who it is. It calls itself EmpalyzeBot and gives an address, so you can write a rule for that one robot.
- It reads robots.txt and obeys it. A page that is disallowed is not read.
- If robots.txt asks for a pause between requests, it waits — up to ten seconds.
- Whatever robots.txt says, at least a second and a half passes between two requests to the same server.
- If the server asks us to slow down, we slow down. A read has both a cap on pages and a deadline, and would rather stop early than sit on your server.
Your rights
You can ask to see what we hold about you, ask us to correct what is wrong, ask for it to be deleted, and ask for it in a format you can take elsewhere. You can also complain to the Norwegian Data Protection Authority (Datatilsynet).
Delete a website from your account and the pages we fetched go with it at once — that is not a request you have to wait on. For everything else, write to privacy email address.
The lawful basis for each category of data must be settled by a lawyer and written in here. We have deliberately not guessed at it.
Changes
If we change what we store or who sees it, we change this page and put a new date at the top. Decide whether registered users should be notified by email of material changes.
September 2026