I’m familiar with HouseFresh, an independent publisher reviewing air purifiers and analyzing indoor air quality to provide unbiased consumer reporting. Danny Ashton founded the site in 2020 and I became aware of the site when they published an article claiming that Google was trying to put independent review sites like theirs out of business in March 2024. There are several videos on my Forensic SEO Youtube channel from that time.
This week a narrative of triumph is spreading across legacy search marketing communities. Prominent industry analysts are pointing to a spectacular resurrection. As one viral post recently broadcasted to the search community:
“I can’t believe my eyes SEOs. HouseFresh, the site that was VERY vocal about Google killing small publishers, has NOW REBOUNDED BY +10,000%!!! Google seems to be actually getting it right again. Sites that actually invest in original content, reviewing the products and putting in the work are getting rewarded.”
The narrative is clean and emotionally satisfying: the independent publisher which became the poster child for Google’s algorithm casualties has finally won its crown back through sheer editorial merit!
There is only one tiny problem with this story.
Prior to their traditional organic search rankings (what we traditionally track as index visibility) rebounding in early June 2026, HouseFresh suffered an eighty percent collapse in real-time AI search citations, how often conversational assistants like Gemini, ChatGPT, and Perplexity reference a site as a source, immediately following the completion of the December 2025 Core Update, according to historical SEMrush data.
A highly curious event.
Most traditional SEO practices assume that if a domain performs well in Google’s organic search rankings, its visibility in AI citations can only improve. Likewise, if a domain is absent from organic SERPs (Search Engine Results Pages), citation rates should decline.
So how is it possible for a domain to drop 80% from its AI citation baseline while over time expanding its legacy organic search rankings?
Seems to defy the laws of SEO gravity.
With that question and while traditional organic rank trackers are celebrating the return of HouseFresh to organic top spots, I performed a forensic audit of the domain’s code-to-visual layout, rendering costs, and AI extraction metrics.
Regrettably, I don’t think we are witnessing a permanent resurrection.
What lies beneath the surface is a different operational reality. It appears we are witnessing the peak of an “Closed-Circuit Bubble” which is driving what, based on cost limits, can only be a temporary state masking a core incompatibility with next-generation search engine architectures – AI.
The GoogleOther Overlay: How the Decoupling Actually Occurred
To understand the divergence between traditional organic search rankings and real-time AI citations, we must analyze how Google decoupled its traditional crawling footprint from indexation and serving.
Daily for almost 5 years (2021-2026) I observed and tested Google’s indexation system, supplemented by server log data.
The two-crawl (or two-wave) indexing system for traditional search did not change. But it did alter silently.
The decoupling was driven by the rollout of the GoogleOther crawler family, engineered to mirror the Googlebot family’s rendering footprints:
- The Core Googlebot Track: The Core Googlebot Track: Traditional organic rankings remain bound to the classic Two-Wave Indexing process. Googlebot Smartphone executes Wave 1 (the HTML-only crawl) and Wave 2 (the expensive Web Rendering Service, or WRS, pass).
- The Decoupled GoogleOther Family: Initiated in April 2023 and completed in November 2025, Google deployed a parallel crawler family to duplicate Googlebot’s footprints for AI training, Retrieval-Augmented Generation (RAG), and developer services:
- April 2023 (Non-Chrome GoogleOther): Google introduces GoogleOther as a non-Chrome agent to handle non-indexation R&D tasks, reducing crawling strain on Googlebot.
- May 2024 (GoogleOther-Video & GoogleOther-Image): Google launches media-specific scrapers to offload asset processing from primary search indexers.
- Mid-2025 Mobile Chrome GoogleOther: Google deploys a mobile-Chrome agent to match Googlebot Smartphone’s rendering footprint.
- November 2025 (Desktop Chrome GoogleOther): Google completes the family, matching Googlebot Desktop’s footprint.
By mirroring Googlebot, Google established a dual-track ecosystem. This directly explains the current HouseFresh “recovery.”
At the time of my recent audit (late June 2026), I couldn’t locate robots.txt file or a sitemap that Googlebot could crawl. I used the Rich Snippet Testing tool to send a mobile or desktop InspectionTool crawler and neither were either available or crawlable.
Likely, if they are live on the server there’s something interfering with Google being able to find and crawl it.
Under RFC 9309 (the Robots Exclusion Protocol), a zero-byte robots file represents a clean slate, instructing crawlers to “allow all.” But for Google’s crawl scheduler—which operates as a resource-conscious cost-accounting system, the impact of a blank robots file and a blank sitemap is profound.
To manage expensive Wave 2 rendering, Google uses a differential delta scheduler to compare <lastmod> sitemap timestamps against its cached index. If no change signals are detected, Googlebot stands down to save GPU compute.
For HouseFresh a 0-byte sitemap, the differential delta scheduler is blinded. It has no sitemap elements to compare and no change-notification triggers. Lacking these signals, Google’s scheduler falls back to an accidental “Structural Freeze” (serving a cached state without triggering a new rendering check).
Think of this like the classic movie “Speed” starring Keanu Reeves and Sandra Bullock. To fool the bad guy watching the bus’s closed-circuit security camera feed, they hijack the signal and set it on a continuous, ten-second loop of the people on the bus. The people are rescued from the bus while the monitor displays absolute stasis of people on the bus.
This is HouseFresh right now.
During Wave 1, Google’s high-speed text parser reads the raw HTML and rewards the deep, qualitative, original product reviews. But because no active sitemap updates trigger the scheduler, the expensive Wave 2 visual rendering and symmetry audits are postponed.
The domain’s massive technical debt is temporarily insulated behind an accidental stasis shield.
(Note: This foundational decoupling model was first isolated and documented in Confessions of an SEO, Season 6. The episodes are available at Confessions of an SEO and streaming on Spotify.)
II. The Systemic 94.4% HTML Code Tax
While traditional search is content to let cached pages ride on their first-pass merits, real-time AI retrieval engines (Gemini, GPTBot, ClaudeBot) do not operate in a stasis vacuum. When a user executes a conversational query, these engines bypass Google’s frozen search cache and instead perform direct, high-velocity HTTP fetches to get the latest live content.
Upon hitting the live pages, they are exposed instantly to a catastrophic 94.4% HTML Code Tax (meaning over 94% of the page’s code is layout bloat rather than actual, readable prose).
To verify the extent of this layout imbalance, I ran audits using the VizzEx Symmetry Gate™ Scanner, a specialized tool designed based on my collaborative research with Kim Albee to measure the mathematical parity between a page’s raw HTML code and its visual layout.
The results reveal a systemic, seemingly platform-wide architectural deficit where code far outweighed the visible content:
Audited URL Topic |
HTML Code Overhead (Tokens) |
Written Prose (Tokens) |
HTML Code Tax (%) |
Gemini/ChatGPT Status |
Perplexity Status |
Claude Status |
|---|---|---|---|---|---|---|
Air Purifiers for Pets |
185,395 |
6,737 |
96.3% |
29% (FAIL) |
51% (FAIL) |
94% (PASS) |
Best Dehumidifiers |
149,685 |
5,665 |
96.3% |
24% (FAIL) |
68% (FAIL) |
94% (PASS) |
Air Purifiers for Mold |
84,280 |
4,479 |
94.9% |
40% (FAIL) |
64% (FAIL) |
100% (PASS) |
Air Purifiers for Dust |
84,688 |
4,963 |
94.4% |
25% (FAIL) |
44% (FAIL) |
94% (PASS) |
Air Purifiers for Allergies |
115,413 |
6,912 |
94.3% |
29% (FAIL) |
48% (FAIL) |
100% (PASS) |
Air Quality Monitors |
112,007 |
7,574 |
93.6% |
19% (FAIL) |
38% (FAIL) |
97% (PASS) |
Air Purifiers for Smoke |
105,918 |
7,411 |
93.4% |
40% (FAIL) |
71% (FAIL) |
94% (PASS) |
Best Air Purifiers We Tested |
172,182 |
14,234 |
91.7% |
29% (FAIL) |
65% (FAIL) |
70% (FAIL) |
Levoit Core 600S Review |
72,789 |
3,963 |
94.5% |
29% (FAIL) |
61% (FAIL) |
100% (PASS) |
Systemic Averages |
120,250 |
6,882 |
94.4% |
FAIL Across All Hubs |
FAIL Across All Hubs |
Lenient Pass (8/9) |
Across all core landing pages, over 94.4% of the page’s token volume is computational noise.
While legacy, offline search indexers can parse this massive bloat over the course of days, real-time RAG (Retrieval-Augmented Generation) models must render and tokenize the DOM dynamically under a strict sub-second execution budget.
Forcing an AI transformer engine to process 94.4% layout noise to extract 5.6% semantic value is a massive, expensive waste of GPU compute (what we call the Compute Tax).
Furthermore, because these pages contain extensive “shadow gaps” (unoptimized structural elements, such as Table of Contents jump-links), the rendering engine must allocate extra memory and processing cycles to resolve the true visual layout. This structural asymmetry triggers automated Fidelity Downgrades (where an AI engine lowers its confidence in the source and refuses to cite it) to protect operating margins, i.e. costs.
III. The Trap of Semantic Absorption & Brand Stripping
This technical debt does not just lock publishers out of citations, it actively facilitates the theft of intellectual property through Semantic Absorption (where the AI learns your facts but completely strips away your brand name).
If a site lacks a hardened, machine-readable signal architecture, AI crawlers will still ingest its raw HTML. Over time, they assimilate those high-quality, hard-won product insights directly into their parametric weights (the AI’s long-term, internal memory), converting your proprietary insights into “common knowledge” that they can output without citation.
The model now “knows” that Model X is the best air purifier for something because someone’s writers spent weeks testing it.
When a conversational searcher asks for a recommendation, the model outputs that exact answer, but it does so without generating a real-time retrieval hop. It won’t cite the domain, send a click, or serve display ads.
Publishers will have essentially donated their proprietary research to train the very systems replacing their traffic.
The Solution: Zero-Friction Geometry
I do not believe this has to be inevitable. To protect intellectual property and survive the transition to real-time generated answers, publishers must adopt an Ingestion Model (re-engineering the site to be highly structured and cheap for AI to read and extract information):
1. Eliminate Fragment Noise
Replace heavy, JS-driven table-of-contents jump-links that trigger immense DOM depth and extraction delays with lightweight, native solutions like the Helpful TOC WordPress plugin to establish true Anchorless Semantic Navigation (clean, lightweight link paths that humans can click and bots can instantly parse).
2. Enforce the Symmetry Gate
Achieve a strict, 1:1 mathematical parity between what your raw HTML source code says to the machine, what your browser-rendered page presents to the human, and what your structured data promises.
3. Deploy VizzEx Pro
By compiling your site under the VizzEx Pro™ server-side static standard , publishers can build an airtight signal architecture. VizzEx Pro implements the VEE (VizzEx Extraction Efficiency) Protocol, wrapping qualitative prose in explicit semantic nodes and hard-coded JSON-LD schema metadata. (Full disclosure, I am a co-founder with Kim Albee at VizzEx and research contributor to this software. It was developed because before now there was no fully capable option to reduce computation costs and connect semantically throughout an entire domain in a way that LLMs can extract information more efficiently.)
By reducing computational noise and formatting content where its Information Gain is explicit to LLMS, this makes citing your brand cheaper than blind ingestion, securing your visibility and traffic inside conversational search.
IV. Scheduler-Level Shelving vs. RAG Retrieval
To understand how a domain can temporarily rank in traditional web search while experiencing an AI citation blackout, we must isolate the crawl scheduler’s resource management from real-time RAG extraction.
- Scheduler-Level Shelving: The crawl scheduler manages URL crawling queues based on Expected Return on Compute (ROC). The crawl scheduler manages queues based on Expected Return on Compute (ROC). When the differential scheduler is blinded, crawling frequency decays. Googlebot stands down, relying on its cached index. The site’s high rendering costs are “shelved” until a major macro-event (like a Core Update) forces a complete, sitewide visual rendering sweep.
- The RAG Retrieval Blockade: The real-time citation engine is entirely separate. Driven by user intent, it executes dynamic, real-time crawls. When it encounters high extraction latency, rendering mismatches, or rate-limiting HTTP 429/503 errors caused by concurrent crawling sweeps, it bypasses the page in favor of low-compute, pre-structured external knowledge hubs.
Server log analysis confirms that “shelving” is an automated resource-management protocol executed at the crawl scheduler level, meaning a site’s crawl budget can be restricted long before its legacy organic rankings reflect the decay.
What we see in this particular case is that without the scheduler getting the message it can let things ride as they are until something triggers the scheduler – drop in AI citations while growing in organic results.
This growth in organic results appears as a “mountain” in Search Console charts. The significance of that is for another writeup, where I’m going to build on the rise of impressions and keyword rankings as what happens when content is ranking solely on the promise of the HTML only.
V. The Crawl Pressure Trap of Legacy Links
A common counter-argument is that current discussions, mentions, and an influx of new editorial backlinks (external links from other sites) will save the domain. That was a satisfactory conclusion pre-LLM, when SEO operated under legacy, document-indexing rules.
Yet here we are in an era of advanced answer environments.
While off-site mentions inform the LLM’s general knowledge weights that “HouseFresh” is an entity associated with “air purifiers,” those mentions on their own do not provide on-page Information Gain (new, extractable facts) to a real-time extraction engine. Instead, these backlinks create a Crawl Pressure Trap:
- When authoritative sites link to HouseFresh, Google’s crawling scheduler is likely forced to increase its crawl frequency to verify the new link pathways.
- Every forced crawl requires Google to repeatedly process the 94.4% HTML code tax.
- Consequently, while the links buy temporary authority, they accelerate the rate at which the domain is flagged as a computational burden in the scheduler’s cost-accounting system.
VI. The Predictive Timeline for Collapse
I believe that HouseFresh’s current organic traffic uptick is a lagging indicator. That it is only a matter of time before their legacy organic search rankings adjust to align with where their AI citations sit today.
While the site’s rich editorial content deserves to rank on its own first-pass merits, this content is delivered through an un-optimized technical architecture that is destined mathematically to be reconciled.
It must be. Economics will determine the timeline.
- The Temporal Buffer: The recent organic recovery, coupled with a steady influx of backlinks and industry mentions, creates a protective “Closed Circuit” that extending their organic search ranking buffer.
- The Settlement: During the pre-audit rendering phases of the upcoming Q3 2026 Core Update, the system likely will reconcile the domain’s high rendering costs with its historical quality metrics.
- The Prediction: Sometime between October 2026 and Q1 2027, the traditional organic index will synchronize with the cost-optimized ingestion models. When this final recalculation occurs, the domain’s traditional organic visibility will align with the AI citation floor established in early 2026.
My Conclusion
A publisher can win in traditional organic search in the short term, but if their content requires too much computational energy to parse, they will lose the AI citation battle. When they lose the AI citation battle, ultimately, they will lose the organic search presence.
Because the ingestion pipelines are now unified, the structural failures that cost a domain its AI citations today will inevitably cost it its canonical rankings tomorrow. The industry may celebrate the “+10,000% rebound” today, but those who understand search engine economics know that the bill always comes due. When that happens is the only unknowable.
Recent Comments