| Metric | Value |
|---|---|
| Domain | recipebook.shofo.ai |
| Category | Computer Vision Training Data |
| Pricing | Unknown |
| Pages Crawled | 6 |
| Crawl Date | 2026-08-04 |
RecipeBook by Shofo Review: Hidden data (52.9/100) — SiteList
RecipeBook by Shofo scores 52.9/100, offering a high-utility marketplace of 25M+ video clips for ML training. While the backlink profile is strong, critical technical blockers like a malformed robots.txt and a lack of metadata render the site nearly invisible to search engines.
Reviewed by SiteList Engine · 34 dimensions · published Examined on August 4, 2026
Quick facts
- Pages Crawled
- 6
- Crawl Date
- 2026-08-04
- Domain
- recipebook.shofo.ai
Executive summary
Shofo.ai is currently in a 'pre-visibility' state. While the underlying product (a massive video dataset) is a high-value linkable asset, the technical implementation of the Next.js frontend and a malformed robots.txt file are preventing search engines from discovering or ranking the site. The site currently lacks a keyword strategy, with metadata that is either too thin or duplicated across functional pages. Fixing these foundational issues is a prerequisite before any content or link-building investment can yield ROI.
Themes
- Indexation Roadblocks: The site is effectively 'dark' to crawlers due to a robots.txt file that returns a 404 error and the total absence of an XML sitemap.
- The Hydration Gap: The 'app-first' architecture hides the marketplace grid behind a JavaScript loading state, leading to a 'Thin Content' profile.
- Identity & Intent Mismatch: Page titles are generic and do not target high-intent keywords like 'Computer Vision Datasets'.
- AI Search Readiness Gap: The site lacks the structured data (Schema.org) required to be cited in AI-driven search results.
- Content Authority Deficit: With zero informational content, the site has no topical authority beyond its brand name.
- Asset Volume
- 25M+ clips
01 · First impressions & positioning — 25M+ clips claimed, but zero customer logos visible
RecipeBook positions itself as a high-utility video data marketplace for ML training, yet it lacks the trust signals required for enterprise adoption. While the claim of 25 million clips is specific and compelling, the homepage provides no proof adjacency, such as customer testimonials or research lab logos. The site relies entirely on its technical utility rather than established authority. This creates a friction point for procurement teams who require evidence of scale and reliability. To improve, the site should anchor its "by the hour" pricing with specific rates and introduce a "Trusted By" section to substantiate its market position.
- Scale Claim
- 25M+ clips
- Social Proof
- Zero logos/testimonials
02 · Audience & messaging — Technical jargon used correctly, but pricing remains opaque
The site uses technical jargon like "FPS" and "Train probe", assuming a high-sophistication audience. However, the question coverage for new users is poor. Critical information regarding hourly rates and data usage rights is missing from the primary crawl, creating high friction for high-intent buyers. While the self-orientation ratio is low—focusing correctly on the user's search task—the lack of "Voice of Customer" evidence remains a significant trust gap. Adding tooltips for technical features and a clear data licensing FAQ would reduce this initial hesitation.
- Technical Terms
- FPS, Train probe, Aspect
- Pricing Transparency
- No rates visible
03 · Usability — Functional search interface hindered by a 3-click contact path
RecipeBook functions well as a technical tool but fails as a self-contained marketplace. The search and filtering mechanics are highly relevant to researchers, allowing for immediate discovery without a login. However, the "business" side of the site is siloed on a separate domain. A user looking for support must click "Our website" in the footer, then navigate to a contact page on shofo.ai. This 3-click journey across domains is a significant friction point for lead generation. Additionally, technical buttons like "Train probe" lack tooltips, which may cause hesitation for first-time visitors unfamiliar with the ecosystem.
- Contact Friction
- 3-click cross-domain path
- Task Prominence
- Immediate search access
04 · Accessibility — Critical labeling failures on /report and no skip links
RecipeBook contains critical accessibility failures for screen reader users, specifically on the /report form. The legal reporting form contains eight inputs with zero accessible names, relying entirely on placeholders that disappear once a user begins typing. This violates WCAG 2.1 AA standards. Furthermore, the site lacks a "Skip to Content" link, forcing keyboard users to tab through the entire sidebar on every page load. The use of Tailwind's "focus:outline-none" also suppresses standard browser focus indicators, making navigation difficult for users with low vision who rely on clear visual cues to track their position.
- Form Accessibility
- 8 inputs missing labels
- Navigation Aids
- hasSkipLink: false
05 · Design execution — Disciplined typography scale vs. 300+ mobile tap-target errors
RecipeBook uses a disciplined design system with only six distinct font sizes and 19 colors. However, this discipline does not extend to mobile ergonomics. Our audit found over 300 instances of interactive elements, such as the "+ Order" and menu buttons, that are significantly smaller than the recommended 44px touch target. Additionally, the primary brand purple (#9f1bfe) fails the 4.5:1 contrast threshold on the site's dark background, computing to 4.02:1. This affects the readability of links and small UI labels for users with visual impairments.
- Tap Target Failures
- 300+ elements < 44px
- Color Contrast
- 4.02:1 (Fail)
07 · Performance — 42MB main domain payload and LCP delayed by client-side rendering
The site is functionally fast once loaded but suffers from structural delays during the initial visit phase. Because the main results grid relies on client-side data fetching, users and search crawlers initially encounter a "Searching..." skeleton state rather than the 25M+ clips promised. This pushes the Largest Contentful Paint (LCP) deep into the timeline. Furthermore, while the subdomain is relatively lean, the parent domain serves a massive 42.1MB payload, indicating a lack of asset-pipeline discipline. We also detected 18 different font faces, which adds unnecessary blocking requests. Transitioning to server-side rendering for the initial results would significantly improve perceived speed.
- Total Payload
- 42.1MB (Main Domain)
- Font Requests
- 18 font faces
09 · Writing quality — 2,822-word legal wall of text and unsupported superlatives
Editorial quality is split between a minimalist app interface and unnavigable legal documentation. The Enterprise Terms page presents over 2,800 words in a single paragraph block, making it nearly impossible to read on any device. Additionally, the site uses unsupported superlatives like "The World's Largest Video Library," which may alienate a technical audience of ML engineers who prefer verifiable metrics. Replacing these claims with the concrete "25M+ clips" figure found in the metadata would build immediate credibility. The site also suffers from duplicate meta descriptions across four functional pages, which hinders search orientation.
- Legal Readability
- 2,822 words in 1 paragraph
- Sentence Complexity
- 49.9 words avg (Legal)
10 · Vertical credibility — Immediate search utility lacks the trust signals of enterprise rivals
RecipeBook adopts a utility-first design register common in developer tools, featuring dark mode and high-density filters. It excels at task prominence, allowing users to search the catalog without the "Book a Demo" walls used by competitors. However, it misses several enterprise-grade conventions found on platforms like Labelbox. In the high-stakes world of AI training data, buyers require evidence of data quality and provenance. The current absence of customer logos, case studies, or quality guarantees creates a credibility gap. Adding a "Data Quality" section explaining how clips are verified would align the site with vertical expectations.
- Trust Signals
- Zero logos/case studies
- Task Access
- Zero-click discovery
11 · Competitive position — Zero indexable content moat compared to Scale AI and Labelbox
RecipeBook is currently invisible to search engines compared to established giants like Scale AI. As a single-page application (SPA) with no XML sitemap and only six crawled pages, the site lacks the content infrastructure required to compete for high-intent keywords like "CV training data." Competitors dominate the SERPs by publishing deep educational content, glossaries, and case studies. RecipeBook has no equivalent informational surface area, meaning it only captures users who already know the brand. Establishing a Wikidata entry and creating a "CV Data Lab" are necessary steps to close this significant authority gap.
- Indexable Pages
- 6 pages
- Entity Authority
- Zero Wikidata results
14 · Authority & link risk — Clean 2025 domain with 25M+ clips as a linkable asset
Shofo.ai is a young domain, registered in February 2025, currently in its authority-building phase. Our audit found no evidence of toxic patterns or expired-domain repurposing. Outbound links are strictly limited to high-trust entities like Hugging Face and official social profiles, which preserves internal equity. The RecipeBook searchable database itself is a significant linkable asset; interactive tools of this scale typically attract high-quality editorial links from the ML research community. The internal structure is efficient, with a maximum click depth of two, ensuring that earned equity flows effectively across the few existing pages.
- Domain Age
- Registered Feb 2025
- Click Depth
- Max 2
15 · Off-page readiness — Strong social signals but missing Organization schema
The brand has a clean social footprint with active profiles on LinkedIn and X, but it lacks the formal entity signals to maximize its off-page potential. No Organization or SoftwareApplication schema was detected, which prevents search engines from explicitly linking the domain to its social entities. While the searchable database is a top-tier linkable asset, the site operates somewhat anonymously without a "Press" or "About" page featuring named founders. Adding a press kit with high-resolution logos and implementing structured data would help Google build the entity graph and improve the site's readiness for earned media mentions.
- Structured Data
- No Schema detected
- Social Probes
- LinkedIn/X (200 OK)
16 · Rank readiness — Duplicate titles on /report cannibalize the homepage brand signal
The site is not currently prepared to rank for category-level keywords due to generic and overlapping metadata. Our crawl found that the /report page shares the exact same title tag as the homepage—"RecipeBook by Shofo"—which causes internal competition for the brand's primary signal. Without unique, keyword-optimized titles, a rank tracker would likely only show visibility for the brand name itself, missing opportunities for "video training data" or "AI video clips." To become rank-ready, each page must be assigned a distinct primary keyword and the malformed robots.txt file must be resolved to ensure consistent indexation.
- Title Duplication
- Shared by 2+ pages
- Keyword Targeting
- Zero non-branded targets
17 · Risk & stability — Feb 2025 domain age and high SERP erosion risk
RecipeBook is a stable but vulnerable new domain with a strong 80/100 score in this dimension. Registered in February 2025, the site lacks legacy baggage but faces significant risk from AI-driven search erosion. Because the marketplace targets technical computer vision data, its core keywords are prime targets for Google's AI Overviews. Without Dataset schema, the site may lose click-through traffic to automated summaries. Establishing niche-relevant backlinks and implementing structured data are the primary paths to stabilizing this new authority.
- Domain Age
- Feb 2025
- SERP Erosion Risk
- High
19 · Editorial QA of content — 66% duplicate metadata and unsupported superlatives
The site's editorial quality is currently weak at 48/100, marked by inconsistent mechanical execution. A primary concern is the use of the unsupported superlative "World's Largest Video Library," which lacks citation or comparative data. Furthermore, 66% of crawled pages feature duplicate titles and meta descriptions, suggesting a templated approach that lacks manual review. The legal pages are particularly dense, with sentence lengths averaging 49 words, significantly exceeding the 20-word target for web clarity. Immediate fixes include differentiating page titles and adding citations for marketplace claims.
- Duplicate Metadata
- 66%
- Average Sentence Length (Legal)
- 49 words
20 · Content program — Zero technical guides for a 25M+ clip marketplace
RecipeBook currently lacks a legible content program, resulting in a critical score of 30/100. For a technical marketplace in the machine learning space, the absence of an engineering hub or data studies is a major competitive disadvantage. There is no informational content to support the research phase where engineers evaluate data quality. Launching a "CV Data Lab" with dataset benchmarks and model optimization guides is essential to build the topical authority required to compete with established rivals. The site must transition from a functional tool to a knowledge resource to earn trust.
- Technical Guides
- 0
- Informational Content URLs
- 0
21 · Distribution & reach — Missing email capture and citation hooks for ML researchers
Distribution is currently a weak point at 40/100, as the site relies on direct traffic without proactive outreach mechanisms. While LinkedIn and X profiles are active, the site lacks a newsletter or RSS feed, which are critical for engaging the ML community. There are no "citation hooks"—such as original data frameworks or technical reports—to earn organic distribution. The current infrastructure is built for one-time transactions rather than the ongoing relationships necessary for a technical brand. Implementing email capture and OpenGraph tags is required to improve social shareability and retention.
- Email Capture Fields
- 0
- RSS/Newsletter Presence
- None
23 · Docs & self-serve help — 404s on all documentation probes for technical users
The site provides no public-facing documentation, resulting in a critical score of 12/100. Probes for common paths like /docs, /api, and /faq all returned 404 errors. This creates significant friction for ML engineers who require technical specifications, data schemas, and licensing details before committing to a purchase. The absence of an llms.txt file further prevents AI-driven research tools from accurately representing the marketplace's 25M+ clips to potential enterprise buyers. Establishing a public help center and a technical spec sheet is a high-priority requirement for self-serve conversion.
- Documentation Path Status
- 404
- AI Discovery File (llms.txt)
- Missing
24 · Measurement readiness — PostHog active via reverse proxy but missing GA4 attribution
Measurement readiness is fair at 62/100, anchored by a sophisticated PostHog implementation. The use of a reverse proxy for analytics suggests technical maturity, but the stack is incomplete. There is no GA4 or Google Tag Manager presence, leaving a significant blind spot in marketing attribution. Furthermore, the site lacks a consent management platform, with the PostHog recorder firing immediately on page load, which poses a compliance risk for global AI audiences. Deploying GTM and a basic consent banner is necessary to align with privacy standards and track traffic sources accurately.
- Marketing Analytics Tags
- 0
- Product Analytics
- PostHog
25 · Technical SEO — Robots.txt returns a 404 HTML page with a noindex tag
Technical SEO is fair at 66/100, but foundational errors hinder crawlability. The robots.txt file is malformed, returning a 404 HTML page that contains a noindex meta tag—a confusing signal for search bots. Additionally, there is no XML sitemap to guide discovery. A significant "negative rendering gap" exists: the raw HTML contains 557 words of data, but the rendered DOM shows only 64 words, suggesting that core value propositions are hidden behind JavaScript hydration. Deploying a plain-text robots.txt and an XML sitemap are the most urgent technical requirements.
- Robots.txt Status
- 404
- Rendered Word Gap
- 493 words
26 · On-page SEO — 52-word rendered count and generic 5-character page titles
On-page SEO is critical at 38/100 due to systemic hygiene failures and thin content. The primary domain title is a generic "Shofo" (5 characters), which fails to signal the site's purpose to search engines. Rendered word counts on core landing pages are as low as 52 words, risking "thin content" flags. The header hierarchy is also flat or missing; the contact page lacks an H1 entirely, and the homepage H1 lacks keyword relevance for "video training data." Rewriting titles to include high-intent keywords and adding 300 words of descriptive copy are essential fixes.
- Rendered Word Count
- 52
- Primary Title Length
- 5 chars
27 · Keyword targeting — Zero non-branded visibility for 'video training data' queries
The site has no discernible keyword strategy beyond brand recognition, scoring 42/100. High-intent terms like "video training data" and "computer vision datasets" are absent from the primary metadata and H1 tags. Currently, the site only ranks for branded terms like "Shofo" and "RecipeBook," which have zero search volume from unacquainted users. To capture commercial intent, the targeting must pivot toward category-level keywords and niche clusters like "hand-object interaction data." Updating the homepage H1 to include "Custom Video Datasets for AI Labs" is a recommended first step.
- Non-Branded Keyword Rank
- 0
- Keyword Relevance (H1)
- Low
29 · Content gaps — Missing 90% of the topical coverage expected for CV data providers
With a critical score of 22/100, the site is effectively invisible to the broader search market. It lacks 90% of the topical coverage expected for a data provider, with zero informational guides or commercial comparison pages. There is no content to intercept users evaluating competitors like Scale AI. The immediate priority is establishing a pillar-and-cluster structure that provides context for the 25M+ clips, moving the site from a "naked" marketplace to an authoritative resource. Creating "Use Case" pages for specific industries like autonomous retail would help bridge this gap.
- Informational URLs
- 0
- Competitor Comparison Pages
- 0
30 · Keyword gaps — Competitors dominate 'video dataset' territory while Shofo remains invisible
Shofo is currently conceding all non-branded search volume to incumbents, resulting in a critical score of 18/100. There is zero overlap in shared territory because the site has not targeted category-level keywords. Competitors like Scale AI and Appen maintain a monopoly on "video dataset" queries. Shofo's unique "by the hour" pricing model is its strongest differentiator, but it remains unindexed and unsearched because it lacks dedicated landing pages targeting affordable training data. Launching a targeted campaign around "High-Quality Video Datasets for ML" is required to contest this territory.
- Shared Keywords with Competitors
- 0
- Category Keyword Presence
- None
32 · AI search readiness — Zero JSON-LD schema and missing llms.txt discovery file
AI search readiness is critical at 29/100 as the site is structurally invisible to AI agents. It lacks all forms of structured data, including Organization and Dataset schema, which are required for citations in AI Overviews. The marketplace is a "black box" to crawlers due to its reliance on client-side rendering. Without an llms.txt file or answer-first informational content, AI engines like Perplexity cannot verify the brand's claims or accurately represent its 25M+ clips. Implementing server-side rendering for search results and deploying a machine-readable llms.txt file are the primary recommendations.
- JSON-LD Schema Blocks
- 0
- AI Accessibility File
- 404
33 · Fix-priority hygiene — 404 robots.txt and hydration gaps blocking indexation
The site's technical foundation is currently compromised by critical indexation roadblocks that prevent search engines from discovering its 25M+ video clips. The most urgent failure is the robots.txt file, which returns a 404 HTML page containing a "noindex" tag, effectively signaling crawlers to ignore the domain. This is compounded by the total absence of an XML sitemap. Furthermore, a severe "hydration gap" exists within the Next.js architecture; while the raw source contains 557 words of data, the rendered view shows only 64 words because the marketplace grid is hidden behind a client-side "Searching..." skeleton state. To restore visibility, the team must deploy a valid plain-text robots.txt, generate a dynamic sitemap, and implement Server-Side Rendering (SSR) to ensure core content is visible to bots. Additionally, unique, keyword-rich title tags must replace the current generic duplicates across the /report and /contact pages.
- robots.txt status
- 404 HTML (noindex)
- Rendered word count
- 64 words
- Raw source word count
- 557 words
- Metadata duplication
- 66%
34 · SEO composite coherence — Technical invisibility concedes 100% of non-branded traffic
RecipeBook by Shofo is a high-potential marketplace currently operating in a state of technical invisibility due to a lack of basic SEO hygiene. While the underlying product—a massive video dataset—is a high-value linkable asset, the site currently concedes all non-branded search traffic to competitors like Scale AI and Labelbox. The coherence of the SEO strategy is undermined by three primary factors: a malformed robots.txt that blocks indexing, a heavy reliance on client-side rendering that creates a "thin content" profile for crawlers, and a total absence of structured data (Schema.org). Without Organization JSON-LD or an llms.txt file, the site is also excluded from AI-driven search results. The 90-day roadmap must prioritize fixing these technical blockers before investing in content. Establishing topical authority will require transitioning from a generic "Shofo" brand signal to targeted keywords like "Computer Vision Training Data" across all metadata and new informational guides.
- AEO score
- 29/100
- Non-branded visibility
- 0%
- Indexable page count
- 6
Verdict — 52.9/100: High-value, hidden data
RecipeBook by Shofo is a high-utility marketplace for AI training data that is currently undermined by its technical execution. The site’s strongest asset is its underlying data—25 million video clips—and a clean, modern design stack. However, it fails on basic SEO hygiene, documentation, and trust signals.
Fixable Weaknesses:
- Technical indexation (robots.txt and sitemap).
- Metadata optimization for high-intent ML keywords.
- Public-facing documentation and licensing details.
Who it's for: ML Engineers and Computer Vision Researchers who already know the brand or are referred directly to the marketplace.
- Overall Score
- 52.9/100
- Authority Score
- 87/100
90-day roadmap
| Window | Action | Modules | Expected effect |
|---|---|---|---|
| Days 1-14 | Fix robots.txt 404 and deploy XML sitemap | Technical, Health | Ensure site discovery and indexation |
| Days 15-45 | Rewrite all Title/H1 tags and implement SSR for the grid | On-page, Performance | Improve keyword relevance and content depth |
| Days 46-90 | Deploy 'Use Case' pages and Organization Schema | Content Gap, AEO | Establish topical authority and AI search citations |
- Roadmap Window
- 90 Days
Methodology & data notes
The review is based on a crawl of 6 pages conducted on 2026-08-04. Several dimensions were excluded: 06, 08, and 18 failed to meet minimum content thresholds, while 12, 13, 22, 28, and 31 were not applicable to this marketplace model. Measurement readiness was assessed via PostHog implementation. For a full explanation of our 34-dimension framework, visit our methodology page.
- Data Sources
- 6-page crawl
Questions buyers actually ask
Is RecipeBook by Shofo a reliable source for ML training data?
The site offers a massive repository of 25M+ video clips, but currently lacks the trust signals, such as case studies or enterprise logos, typically expected by professional ML researchers.
Why is RecipeBook not appearing in search results?
Our crawl identified critical technical issues, including a malformed robots.txt file and a lack of unique metadata, which prevent search engines from indexing the marketplace's content.
Does the site support self-service licensing?
No. At the time of review, there is no public-facing documentation or self-serve help center explaining licensing terms or technical specifications for the datasets.
How does RecipeBook compare to competitors like Scale AI?
While the core data asset is significant, RecipeBook lacks the content infrastructure and keyword strategy that established competitors use to dominate the computer vision market.