For years, SEOs have speculated about the inner workings of Google’s search algorithm. Now, a leak of internal Google API documents offers findings may surprise the casual user, but many of them validate long-held SEO beliefs. For me it is interesting to see which things do not get much coverage for example, spam-links and how they are ignored or even punished, near content type (plagiarism), structured data errors and generic web performance. Finally conversions and goal measuring.

Things which do not get any coverage include recent ‘real life’ WebVitals data to reclassify a page.

Other aspects I find interesting is the absence of ‘web page speed‘ and ‘mobile first‘. This suggests the likelihood that this ranking API is one piece of the puzzle, the data-hole is indeed very deep.

No reference to ‘disavow‘ or ‘spam‘ teams. Ultimately ranking seems to act out of user ‘happiness’ and clicks, links, and if you are lucky the Google preference bias (which should be another case for an anti-trust trial).

One key takeaway is the significant role user engagement plays in search rankings. The leaked documents detail a system called “NavBoost” that analyses clicks, categorizing them as positive (“good clicks”), negative (“bad clicks”), or even super valuable (“unicorn clicks”). Metrics like dwell time, click-through rate, and how often users refine searches after visiting a site all factor in. This reinforces the importance of creating content that keeps users engaged and satisfied.

Another interesting revelation is the potential use of Chrome browsing data for ranking purposes. The documents mention metrics like “uniqueChromeViews,” suggesting Google must leverage Chrome user behaviour data to understand how users interact with websites. This could explain the growing emphasis on factors like website speed and mobile-friendliness.

I hold that Chrome was likely created to gather clickstream data direct from the user and, now users are prompted to log-into Chrome browser your own identification and authority is likely held in big data too.

The leak also sheds light on the structure of Google’s ranking system. It appears to be a multi-layered approach, not a single monolithic algorithm and it has patches on top of patches, what they call filters. But we know this anyway. I have said to clients there is a big update with smaller payloads as I see effects for localised pages being hit, when it was supposed just to affect affiliate sites. Also, the filters explains how there are various rollbacks with some particular hard hitting algorithm updates.

Various microservices, like the well-known Panda, work together as filters, each focusing on specific ranking factors like backlinks, content quality, and user signals. This reinforces the need for a well-rounded SEO strategy that considers all aspects of search engine optimisation and if you have industry knowledge of historic updates, it gives you an insight how Google keep on adding filters on filters to constantly evolve ‘rank factors’.

Finally, the documents hint at a system akin to “domain authority,” with metrics like “siteAuthority” and “authorityPromotion.” This, along with mentions of “PageRank” and “linkInfo,” suggests that backlinks remain a crucial ranking factor. However, Google seems to be keeping an eye on unnatural link growth with metrics like “PhraseAnchorSpamDays.” This emphasises the importance of building high-quality backlinks over time, and not using manipulative growth tactics.

Overall, the Google API leak provides valuable insights into the complex world of search engine ranking. But for me there are big holes with let’s say, Google directives for mobile first, page speed and technical updates which suggest another factor. While the information should be interpreted with some caution, it offers a strong confirmation of many SEO best practices. 

Here are the main things that have come out of the leaked ranking score API and how it effects SEO professional knowledge and websites.

Human Quality Raters Have a Influence on Rankings

Attributes like “furballUrl” and “humanRatings” indicate data from Google’s quality raters may impact rankings. Having monitored their visits over websites and watching what pages they visit it is quite interesting.

Domain Authority Metric

The leak revealed metrics like “siteAuthority” and “authorityPromotion”, suggesting Google may use a domain authority-like metric for ranking, contrary to their denials.

The Importance Of Click Data

Systems like “NavBoost” and “Glue” highlight Google’s extensive use of click data and user engagement signals like clicks, dwell time etc. as major ranking factors.

Use of Chrome Browser Data

The leak reveals Google tracks user interactions and clickstream data from the Chrome browser, contradicting previous claims that Chrome data is not used for ranking purposes. Metrics like “uniqueChromeViews” show Google collects site-level Chrome usage signals. This was suggested back in 2013 from Matt Cutts and perhaps the reason why Chrome was developed was to get clickstream data :o. Metrics like “uniqueChromeViews” and “chromeInTotal” suggest Google collects site-level Chrome usage data to inform its ranking systems. 

Click Data and User Engagement Signals

Contrary to Google’s public statements, the documents expose extensive use of click data and user engagement metrics as major ranking factors. Systems like “NavBoost” leverage clickstream data including clicks, dwell time, query reformulations etc. to rank pages. In essence, NavBoost is a ‘relevance’, ‘user experience’ and ‘intent’ based ranking formula. For those saying users that ‘pogo’ on and off pages and ranks – this discovery goes out to you!

NavBoost allows Google to closely monitor how users interact with search results and websites, using that data as a strong ranking factor to prioritise the most satisfying results for users based on observed patterns and engagement.

Key aspects of NavBoost:

  1. Click Types
  • It categorizes clicks into different types like “good clicks”, “bad clicks”, “last longest clicks” etc. to identify positive vs negative user interactions.
  • “Unicorn clicks” seem to refer to highly valuable click patterns.
  • It distinguishes between “squashed” (devalued) and “unsquashed” (valued) clicks to filter out spam.
  1. User Signals
  • NavBoost heavily relies on click data like number of clicks, click-through rates, dwell time (last longest clicks) etc. as ranking signals.
  • It tracks impressions to measure how often a result is shown.
  • It likely utilizes other engagement signals like hovers, scrolls etc. through the “Glue” system.
  1. Data Processing
  • It maintains a 13-month rolling window of user data for analysis.
  • Click data is segmented based on factors like location, device type etc. through “slices”.
  • It scores queries based on identified user intent (informational, navigational etc.)
  1. Ranking Impact
  • NavBoost evaluates and scores websites/pages based on the collected user signals.
  • This scoring can result in boosting or demoting of rankings at the URL, subdomain or domain level.
  • It acts as a re-ranking system, adjusting initial rankings based on its user-focused analysis.
  1. UX and CTR Optimisation
  • Having a user-friendly website with clear navigation seems to be rewarded by NavBoost signals.
  • Factors like brand reputation and optimising for SERP features may help attract more valuable clicks.

Who Are The Unicorn Users

The leaked documents mention “unicorn clicks” which appear to refer to specific types of clicks or user interactions that Google considers highly valuable or significant. 

While the exact definition is unclear, it likely relates to identifying high-quality user engagement signals that could positively impact rankings (see NavBoost)

Google Sandbox For New Websites Exists

Despite Google’s denials, the leaked documents confirm the existence of the “Google Sandbox” – a filter that prevents new websites from ranking well initially to combat spam. 

This validates what many SEOs and myself have suspected through testing with new websites and trying to get them to rank or even indexed correctly.

Google’s Bias Exists: Whitelists for Certain Verticals

The leak confirms Google maintains “whitelists” that restrict certain types of results for verticals like travel, COVID-19 information, elections and more, likely for quality and safety reasons. Metrics like “isElectionAuthority” and “isCovidLocalAuthority” show Google maintains whitelists restricting certain types of results.

This also demonstrates that Google could give preferential treatment to specific websites if it wanted to. You mean some big brand websites which do a lot of advertising and their SEO is terrible. Yes, I think this answers that concern, I mean conspiracy theory ;D

Algorithm Updates are Filters as Layered Ranking Systems

Rather than a single algorithm, Google’s ranking process involves multiple microservices or systems that act as successive filters, each adjusting rankings based on specialised factors like links, content quality, user signals etc. There’s over 14,000 ranking factors. Most which we know already are the ten principle areas of EEAT, with let’s say, lots of layers, like an onion over the top of each.

Each multiple micro-services or algorithms that act as filters, each handling different factors. The results get refined through successive layers, with various systems like NavBoost, Panda, Twiddlers etc. adjusting the rankings based on their specialised purpose.

Panda Scoring System Is Current and is Important as Ever

Details emerge on how the influential Panda algorithm likely operates – by assigning websites a “Site Quality Score” based on UX, content quality and backlink profile, which then boosts or demotes the entire site’s rankings.

I would assume other large algorithm updates, the historical ones Penguin, Possum add into this and exist as filters too.

Re-directs and Link Juice Passing uses Historical Page Versions

The “urlHistory” system stores around 20 previous versions of pages, impacting redirects and link equity flow. It suggests that Google keeps versions of pages to look at link equity and the type of old to new content. I have always expected this as 301’s and other redirects never really seem to give what they once did.

Ranking Demotions – Experience and UX

Various “demotion” factors like poor UX, irrelevant locations, low-quality reviews etc. can significantly lower rankings. This goes under user experience and experience of EEAT.

Spammy Anchor Text Relevance: Use it with Care

Metrics related to anchor text relevance indicate mismatches between anchor text and target content can demote sites and rankings. Importance of anchor text relevance and potential “demotion” for mismatches looks very strong.

The anchor text must accurately describe the content of the target site or web page. Getting it wrong may actually get demoted in the search results.

Google Caffeine Is Ongoing – Content Freshness

Multiple metrics like “contentAge” and “lastSignificantUpdate” allow Google to evaluate and prioritise fresh, updated content. So yes, Caffeine is a part of the ranking – something I have harped on about for years.

Content Length and Structured Content – Inverted Pyramid

Metrics like “numTokens” suggest Google has limits on content processing length, prioritising important text upfront. I rarely if ever see this on some websites ranking well, suggesting there is a lot of biasing plus sandboxing going on. However, using inverted pyramid I have always seen positive results.

Unicorn Biased Author Material

The “isAuthor” function and “authorName” suggest Google may use authorship as a ranking signal. This suggests another positive bias and shield to entry for newer and upcoming, but none the less, brilliant authors with great insights.

Like You Didn’t Know: Links Are Important

Frequent mentions of “PageRank”, “linkInfo” and links as modifiers reinforce the continued importance of high-quality backlinks, or better said, categorisation of links into quality tiers based on source. We know that PageRank was a retired category it is interesting to see that it still exists in recent memory for Google.

I think it’s great to add here that links and the traffic they send is important. This is often overlooked.

Quickly Scaling Links Gets Noticed

Once an SEO myth to some that quickly getting links, gets the all seeing eye of Google on you is actually true. Metrics like “PhraseAnchorSpamDays” suggest Google monitors backlink velocity to identify unnatural link growth.

Domain Registration Data

Google likely tracks domain ownership and registration details, though the impact on rankings is unclear although we have seen expired domain abuse a part of algorithmic updates in March 2024. This links into Geo-fencing signals based on location. Also, if you are setting up various spam sites to favour each other, the favour I would imagine is nerfed!

Conclusion: Google Leaked Ranking API & SEO Takeaways

While many details have long been speculated, this leak confirms and quantifies the significant role various factors play in determining rankings. There remains a lot of unknowns and how deep data fits into this, especially with mobile, schema and page speed.

User signals and engagement metrics like clicks, dwell time, and query reformulations are heavily weighted, with systems like “NavBoost” closely monitoring these “unicorn clicks” to boost results. 

Google also leverages data from the Chrome browser, contradicting previous denials, to gauge site-level usage patterns.

Rather than a single algorithm, Google employs a layered, multi-algorithm approach with specialized systems like Panda, Twiddlers, and others refining the results based on factors like links, content quality, and user signals.

Panda, in particular, assigns sites a “Quality Score” based on UX, content, and backlinks to demote or promote entire domains.

Metrics like “siteAuthority” and “authorityPromotion” suggest Google utilizes a domain authority-like metric, while frequent mentions of “PageRank” and “linkInfo” reaffirm the enduring value of high-quality backlinks.

Google also monitors backlink velocity to identify unnatural link growth patterns.

The leak sheds light on other key areas like content freshness scoring, processing limits on content length, the impact of human quality raters, and the use of whitelists for certain verticals.

It confirms long-held suspicions about Google’s “Sandbox” filter for new sites and the importance of factors like anchor text relevance.

Overall, the leak validates many SEO principles while revealing new insights into the complex, multi-layered algorithms. I found what was not revealed to be perplexing. As should you!

A. Leaton
Latest posts by A. Leaton (see all)

Published by A. Leaton

Over 12 years’ experience in diverse digital marketing settings within international start-up and corporate environments. Proven ability to deliver conversion and traffic uplift results through hands-on execution of multilingual and international SEO, email, content, social media, PPC, broadsheet PR, web-development, email marketing, and CRM strategies and campaigns.

Leave a comment

Your email address will not be published. Required fields are marked *