In short: Google had indexed only 89 of my 156 pages (57%), even though the technical side was clean. The real cause was 54 portfolio pages with forty to sixty words each eating my crawl budget. After rewriting them properly, indexing rose to 99 pages (63%).
Google is not indexing your pages. That is the blunt version of what I suspected about my own site for a long time, without ever measuring it. I mean whether a page exists in Google's database at all, not how it ranks. If it does not, the keyword, the copy and the design are all irrelevant.
I have now queried every one of my 156 URLs individually, through Search Console's own API. The result: 89 indexed, that is 57%. Every third page of mine did not exist as far as Google was concerned.
This article is about what the cause turned out to be. It was not what I expected, and not what most articles point at either.
If Google is not indexing your pages, rule this out first
The logical first suspect is the technical side. I went through the non-indexed pages one by one:
- HTTP 200, all reachable
- self-referencing canonical, no mis-canonicalisation
index, followin the robots meta, nothing excluded- correct hreflang triad between the Hungarian, Romanian and English versions
- internal links: the English services page gets three, the about page four, the contact page five, from the English homepage alone
So: nothing. Zero technical faults. Reassuring the first time, uncomfortable the second: if the technical side is fine, then the problem is something a setting will not fix.
The real cause: 54 pages with forty words each
When I broke the index list down by page type, the point emerged. Of the 67 missing URLs, 37 were my portfolio pages.
I measured why. The body text of the reference pages:
- Rubikom: 34 words
- Evolut Agency: 43 words
- Mediner: 59 words
- RegeNail: 66 words
For comparison, one of my service pages is 230 words. And I had 18 projects in three languages, so that is 54 URLs with forty to sixty words each.
Google does not index a forty-word page. That is understandable: there is nothing on it to index. The uncomfortable part is that crawling them still costs budget.
What crawl budget is, plainly
Google does not walk the web with unlimited resources. Every domain has an approximate budget for how much it is willing to crawl there. According to Google's own documentation on crawl budget, that budget grows with the domain's authority. My site has a Domain Rating of 7 according to Ahrefs, which is very low. The Romanian average is around 35.
So: a small budget, and I spent 54 URLs of it on pages that had forty words on them.
This mistake is easy to make without noticing, because every individual decision looked sensible. Give every project its own page: sensible. Do it in all three languages: sensible. Only together did it become a set of 54 items siphoning budget away from the pages that earn money.
What I did
There were two routes. One: set the reference pages to noindex and leave only the collecting archive indexable. That reallocates the budget immediately and is half an hour of work.
I chose the other: I wrote all of them properly. Every reference got a 200-280 word case study (what the brief was, what I did, what came of it) in all three languages.
Two things I deliberately left out.
I did not write numbers as results. There is not a single "30% more enquiries", because I cannot verify it. On a reference page an invented number is the worst possible thing: it takes away exactly the credibility the page exists for.
On partnership projects I stated what is not my work. On several references I was not the sole developer but a partner in an agency team, and there the design and the strategy are not to my credit. Those pages got a separate section saying so, because a case study claiming more than is true is a verifiable lie.
Alongside that I requested indexing for eleven pages in Search Console. Indexing went from 89 to 99, from 57% to 63%, within hours.
What you would try in vain
Along the way I also ran into three things that get sold as solutions and are not.
The Google Indexing API. It exists, it works, and it is exclusively for JobPosting and BroadcastEvent content, meaning job postings and live broadcasts. Submitting an ordinary subpage through it is against the rules and ineffective. If someone offers this as "instant indexing", they either do not know what they are doing, or they do.
IndexNow towards Google. IndexNow is a fine standard: you submit that a URL has changed and the engine takes a look. Bing, Yandex, Seznam and Naver use it. Google does not.
Repeating the "request indexing" click. The daily quota is for new and updated pages. Requesting the same unchanged URL every day speeds up nothing, and Google's own documentation says so. It is worth going through the missing pages once, not through the same ones ten times.
The uncomfortable discovery: Bing and ChatGPT
Since I was already at IndexNow, I checked my own setup. The key file was there in the web root, properly. And nothing was calling the API. For months I believed it was configured, while it did nothing at all.
That stung more than it should have, because Bing today is not the periphery many take it for. ChatGPT's search and Microsoft Copilot are built on the Bing index. What Bing does not know, no language model can cite.
I checked what Bing knows about me: 26 indexed URLs out of 156. And the sitemap had not been submitted, which was Bing Webmaster Tools' own first recommendation too. The Bing AI Performance report, which shows how often Copilot cited the site: zero.
I wrote the missing automation, submitted the sitemap, and the initial push went out with 156 URLs. One thing I checked in advance, though, because it is not a given: whether my server lets the AI crawlers through at all. GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, bingbot and Applebot all get a 200 response and full HTML. That is not the case everywhere: the same host rejects some other testing tools with a 403.
What I would check on your site
Three steps, in order, and none of them needs a paid tool.
1. Look at the real number. Search Console, Pages report: how many URLs are indexed and how many are not. If your sitemap has a hundred URLs and thirty are indexed, your problem is not the keywords.
2. Count the thin pages. If you have twenty or thirty pages with body text under a hundred words, those are eating your budget. Either write them properly or take them out of the index. Just do not leave them sitting there in that state.
3. Check whether Bing knows about you. It is free, it can be imported from Search Console, and from then on you can see what the index ChatGPT's search is built on knows about you.
And the part at the end
Everything I have described here is hygiene. It is necessary and worth doing, but it is not the bottleneck.
The bottleneck is that barely any other site links to my domain. Google says so, Ahrefs' Domain Rating of 7 says so, and Bing says it in its own words too: "your site does not have enough inbound links from high quality domains". Four independent sources pointing at the same single thing.
The crawl budget was tight because the authority is low. Indexing is slow because the authority is low. That cannot be fixed with a setting, and it cannot be bought in link packages either. My own measurement shows it: I have 401 links from 346 domains, and the Domain Rating is still 7. Quantity is worth nothing.
What is worth something: genuine, justifiable mentions from places that actually relate to you. That is what I am working on now, and I will write about what it brought too.
If you are not sure whether what you published on your own site is indexed at all, send me a message and I will take a look. And if you are interested in what counts as a real issue and what is noise in an SEO audit, I wrote a separate article about that.

