Skip to content
SiteList

Partial examination: some dimensions could not be assessed in this inspection. Missing dimensions are excluded from the score and listed in the methodology notes below.

RecipeBook by Shofo Review: Hidden data (52.9/100) — SiteList

RecipeBook by Shofo scores 52.9/100, offering a high-utility marketplace of 25M+ video clips for ML training. While the backlink profile is strong, critical technical blockers like a malformed robots.txt and a lack of metadata render the site nearly invisible to search engines.

Reviewed by SiteList Engine · 34 dimensions · published Examined on August 4, 2026

Quick facts

Metric Value
Domain recipebook.shofo.ai
Category Computer Vision Training Data
Pricing Unknown
Pages Crawled 6
Crawl Date 2026-08-04
Evidence
Pages Crawled
6
Crawl Date
2026-08-04
Domain
recipebook.shofo.ai

Executive summary

Shofo.ai is currently in a 'pre-visibility' state. While the underlying product (a massive video dataset) is a high-value linkable asset, the technical implementation of the Next.js frontend and a malformed robots.txt file are preventing search engines from discovering or ranking the site. The site currently lacks a keyword strategy, with metadata that is either too thin or duplicated across functional pages. Fixing these foundational issues is a prerequisite before any content or link-building investment can yield ROI.

Themes

  • Indexation Roadblocks: The site is effectively 'dark' to crawlers due to a robots.txt file that returns a 404 error and the total absence of an XML sitemap.
  • The Hydration Gap: The 'app-first' architecture hides the marketplace grid behind a JavaScript loading state, leading to a 'Thin Content' profile.
  • Identity & Intent Mismatch: Page titles are generic and do not target high-intent keywords like 'Computer Vision Datasets'.
  • AI Search Readiness Gap: The site lacks the structured data (Schema.org) required to be cited in AI-driven search results.
  • Content Authority Deficit: With zero informational content, the site has no topical authority beyond its brand name.
Evidence
Asset Volume
25M+ clips

01 · First impressions & positioning — 25M+ clips claimed, but zero customer logos visible

RecipeBook positions itself as a high-utility video data marketplace for ML training, yet it lacks the trust signals required for enterprise adoption. While the claim of 25 million clips is specific and compelling, the homepage provides no proof adjacency, such as customer testimonials or research lab logos. The site relies entirely on its technical utility rather than established authority. This creates a friction point for procurement teams who require evidence of scale and reliability. To improve, the site should anchor its "by the hour" pricing with specific rates and introduce a "Trusted By" section to substantiate its market position.

Evidence
Scale Claim
25M+ clips
Social Proof
Zero logos/testimonials

02 · Audience & messaging — Technical jargon used correctly, but pricing remains opaque

The site uses technical jargon like "FPS" and "Train probe", assuming a high-sophistication audience. However, the question coverage for new users is poor. Critical information regarding hourly rates and data usage rights is missing from the primary crawl, creating high friction for high-intent buyers. While the self-orientation ratio is low—focusing correctly on the user's search task—the lack of "Voice of Customer" evidence remains a significant trust gap. Adding tooltips for technical features and a clear data licensing FAQ would reduce this initial hesitation.

Evidence
Technical Terms
FPS, Train probe, Aspect
Pricing Transparency
No rates visible

03 · Usability — Functional search interface hindered by a 3-click contact path

RecipeBook functions well as a technical tool but fails as a self-contained marketplace. The search and filtering mechanics are highly relevant to researchers, allowing for immediate discovery without a login. However, the "business" side of the site is siloed on a separate domain. A user looking for support must click "Our website" in the footer, then navigate to a contact page on shofo.ai. This 3-click journey across domains is a significant friction point for lead generation. Additionally, technical buttons like "Train probe" lack tooltips, which may cause hesitation for first-time visitors unfamiliar with the ecosystem.

Evidence
Contact Friction
3-click cross-domain path
Task Prominence
Immediate search access

04 · Accessibility — Critical labeling failures on /report and no skip links

RecipeBook contains critical accessibility failures for screen reader users, specifically on the /report form. The legal reporting form contains eight inputs with zero accessible names, relying entirely on placeholders that disappear once a user begins typing. This violates WCAG 2.1 AA standards. Furthermore, the site lacks a "Skip to Content" link, forcing keyboard users to tab through the entire sidebar on every page load. The use of Tailwind's "focus:outline-none" also suppresses standard browser focus indicators, making navigation difficult for users with low vision who rely on clear visual cues to track their position.

Evidence
Form Accessibility
8 inputs missing labels
Navigation Aids
hasSkipLink: false

05 · Design execution — Disciplined typography scale vs. 300+ mobile tap-target errors

RecipeBook uses a disciplined design system with only six distinct font sizes and 19 colors. However, this discipline does not extend to mobile ergonomics. Our audit found over 300 instances of interactive elements, such as the "+ Order" and menu buttons, that are significantly smaller than the recommended 44px touch target. Additionally, the primary brand purple (#9f1bfe) fails the 4.5:1 contrast threshold on the site's dark background, computing to 4.02:1. This affects the readability of links and small UI labels for users with visual impairments.

Evidence
Tap Target Failures
300+ elements < 44px
Color Contrast
4.02:1 (Fail)

07 · Performance — 42MB main domain payload and LCP delayed by client-side rendering

The site is functionally fast once loaded but suffers from structural delays during the initial visit phase. Because the main results grid relies on client-side data fetching, users and search crawlers initially encounter a "Searching..." skeleton state rather than the 25M+ clips promised. This pushes the Largest Contentful Paint (LCP) deep into the timeline. Furthermore, while the subdomain is relatively lean, the parent domain serves a massive 42.1MB payload, indicating a lack of asset-pipeline discipline. We also detected 18 different font faces, which adds unnecessary blocking requests. Transitioning to server-side rendering for the initial results would significantly improve perceived speed.

Evidence
Total Payload
42.1MB (Main Domain)
Font Requests
18 font faces

09 · Writing quality — 2,822-word legal wall of text and unsupported superlatives

Editorial quality is split between a minimalist app interface and unnavigable legal documentation. The Enterprise Terms page presents over 2,800 words in a single paragraph block, making it nearly impossible to read on any device. Additionally, the site uses unsupported superlatives like "The World's Largest Video Library," which may alienate a technical audience of ML engineers who prefer verifiable metrics. Replacing these claims with the concrete "25M+ clips" figure found in the metadata would build immediate credibility. The site also suffers from duplicate meta descriptions across four functional pages, which hinders search orientation.

Evidence
Legal Readability
2,822 words in 1 paragraph
Sentence Complexity
49.9 words avg (Legal)

10 · Vertical credibility — Immediate search utility lacks the trust signals of enterprise rivals

RecipeBook adopts a utility-first design register common in developer tools, featuring dark mode and high-density filters. It excels at task prominence, allowing users to search the catalog without the "Book a Demo" walls used by competitors. However, it misses several enterprise-grade conventions found on platforms like Labelbox. In the high-stakes world of AI training data, buyers require evidence of data quality and provenance. The current absence of customer logos, case studies, or quality guarantees creates a credibility gap. Adding a "Data Quality" section explaining how clips are verified would align the site with vertical expectations.

Evidence
Trust Signals
Zero logos/case studies
Task Access
Zero-click discovery

11 · Competitive position — Zero indexable content moat compared to Scale AI and Labelbox

RecipeBook is currently invisible to search engines compared to established giants like Scale AI. As a single-page application (SPA) with no XML sitemap and only six crawled pages, the site lacks the content infrastructure required to compete for high-intent keywords like "CV training data." Competitors dominate the SERPs by publishing deep educational content, glossaries, and case studies. RecipeBook has no equivalent informational surface area, meaning it only captures users who already know the brand. Establishing a Wikidata entry and creating a "CV Data Lab" are necessary steps to close this significant authority gap.

Evidence
Indexable Pages
6 pages
Entity Authority
Zero Wikidata results

14 · Authority & link risk — Clean 2025 domain with 25M+ clips as a linkable asset

Shofo.ai is a young domain, registered in February 2025, currently in its authority-building phase. Our audit found no evidence of toxic patterns or expired-domain repurposing. Outbound links are strictly limited to high-trust entities like Hugging Face and official social profiles, which preserves internal equity. The RecipeBook searchable database itself is a significant linkable asset; interactive tools of this scale typically attract high-quality editorial links from the ML research community. The internal structure is efficient, with a maximum click depth of two, ensuring that earned equity flows effectively across the few existing pages.

Evidence
Domain Age
Registered Feb 2025
Click Depth
Max 2

15 · Off-page readiness — Strong social signals but missing Organization schema

The brand has a clean social footprint with active profiles on LinkedIn and X, but it lacks the formal entity signals to maximize its off-page potential. No Organization or SoftwareApplication schema was detected, which prevents search engines from explicitly linking the domain to its social entities. While the searchable database is a top-tier linkable asset, the site operates somewhat anonymously without a "Press" or "About" page featuring named founders. Adding a press kit with high-resolution logos and implementing structured data would help Google build the entity graph and improve the site's readiness for earned media mentions.

Evidence
Structured Data
No Schema detected
Social Probes
LinkedIn/X (200 OK)

16 · Rank readiness — Duplicate titles on /report cannibalize the homepage brand signal

The site is not currently prepared to rank for category-level keywords due to generic and overlapping metadata. Our crawl found that the /report page shares the exact same title tag as the homepage—"RecipeBook by Shofo"—which causes internal competition for the brand's primary signal. Without unique, keyword-optimized titles, a rank tracker would likely only show visibility for the brand name itself, missing opportunities for "video training data" or "AI video clips." To become rank-ready, each page must be assigned a distinct primary keyword and the malformed robots.txt file must be resolved to ensure consistent indexation.

Evidence
Title Duplication
Shared by 2+ pages
Keyword Targeting
Zero non-branded targets

17 · Risk & stability — Feb 2025 domain age and high SERP erosion risk

RecipeBook is a stable but vulnerable new domain with a strong 80/100 score in this dimension. Registered in February 2025, the site lacks legacy baggage but faces significant risk from AI-driven search erosion. Because the marketplace targets technical computer vision data, its core keywords are prime targets for Google's AI Overviews. Without Dataset schema, the site may lose click-through traffic to automated summaries. Establishing niche-relevant backlinks and implementing structured data are the primary paths to stabilizing this new authority.

Evidence
Domain Age
Feb 2025
SERP Erosion Risk
High

19 · Editorial QA of content — 66% duplicate metadata and unsupported superlatives

The site's editorial quality is currently weak at 48/100, marked by inconsistent mechanical execution. A primary concern is the use of the unsupported superlative "World's Largest Video Library," which lacks citation or comparative data. Furthermore, 66% of crawled pages feature duplicate titles and meta descriptions, suggesting a templated approach that lacks manual review. The legal pages are particularly dense, with sentence lengths averaging 49 words, significantly exceeding the 20-word target for web clarity. Immediate fixes include differentiating page titles and adding citations for marketplace claims.

Evidence
Duplicate Metadata
66%
Average Sentence Length (Legal)
49 words

20 · Content program — Zero technical guides for a 25M+ clip marketplace

RecipeBook currently lacks a legible content program, resulting in a critical score of 30/100. For a technical marketplace in the machine learning space, the absence of an engineering hub or data studies is a major competitive disadvantage. There is no informational content to support the research phase where engineers evaluate data quality. Launching a "CV Data Lab" with dataset benchmarks and model optimization guides is essential to build the topical authority required to compete with established rivals. The site must transition from a functional tool to a knowledge resource to earn trust.

Evidence
Technical Guides
0
Informational Content URLs
0

21 · Distribution & reach — Missing email capture and citation hooks for ML researchers

Distribution is currently a weak point at 40/100, as the site relies on direct traffic without proactive outreach mechanisms. While LinkedIn and X profiles are active, the site lacks a newsletter or RSS feed, which are critical for engaging the ML community. There are no "citation hooks"—such as original data frameworks or technical reports—to earn organic distribution. The current infrastructure is built for one-time transactions rather than the ongoing relationships necessary for a technical brand. Implementing email capture and OpenGraph tags is required to improve social shareability and retention.

Evidence
Email Capture Fields
0
RSS/Newsletter Presence
None

23 · Docs & self-serve help — 404s on all documentation probes for technical users

The site provides no public-facing documentation, resulting in a critical score of 12/100. Probes for common paths like /docs, /api, and /faq all returned 404 errors. This creates significant friction for ML engineers who require technical specifications, data schemas, and licensing details before committing to a purchase. The absence of an llms.txt file further prevents AI-driven research tools from accurately representing the marketplace's 25M+ clips to potential enterprise buyers. Establishing a public help center and a technical spec sheet is a high-priority requirement for self-serve conversion.

Evidence
Documentation Path Status
404
AI Discovery File (llms.txt)
Missing

24 · Measurement readiness — PostHog active via reverse proxy but missing GA4 attribution

Measurement readiness is fair at 62/100, anchored by a sophisticated PostHog implementation. The use of a reverse proxy for analytics suggests technical maturity, but the stack is incomplete. There is no GA4 or Google Tag Manager presence, leaving a significant blind spot in marketing attribution. Furthermore, the site lacks a consent management platform, with the PostHog recorder firing immediately on page load, which poses a compliance risk for global AI audiences. Deploying GTM and a basic consent banner is necessary to align with privacy standards and track traffic sources accurately.

Evidence
Marketing Analytics Tags
0
Product Analytics
PostHog

25 · Technical SEO — Robots.txt returns a 404 HTML page with a noindex tag

Technical SEO is fair at 66/100, but foundational errors hinder crawlability. The robots.txt file is malformed, returning a 404 HTML page that contains a noindex meta tag—a confusing signal for search bots. Additionally, there is no XML sitemap to guide discovery. A significant "negative rendering gap" exists: the raw HTML contains 557 words of data, but the rendered DOM shows only 64 words, suggesting that core value propositions are hidden behind JavaScript hydration. Deploying a plain-text robots.txt and an XML sitemap are the most urgent technical requirements.

Evidence
Robots.txt Status
404
Rendered Word Gap
493 words

26 · On-page SEO — 52-word rendered count and generic 5-character page titles

On-page SEO is critical at 38/100 due to systemic hygiene failures and thin content. The primary domain title is a generic "Shofo" (5 characters), which fails to signal the site's purpose to search engines. Rendered word counts on core landing pages are as low as 52 words, risking "thin content" flags. The header hierarchy is also flat or missing; the contact page lacks an H1 entirely, and the homepage H1 lacks keyword relevance for "video training data." Rewriting titles to include high-intent keywords and adding 300 words of descriptive copy are essential fixes.

Evidence
Rendered Word Count
52
Primary Title Length
5 chars

27 · Keyword targeting — Zero non-branded visibility for 'video training data' queries

The site has no discernible keyword strategy beyond brand recognition, scoring 42/100. High-intent terms like "video training data" and "computer vision datasets" are absent from the primary metadata and H1 tags. Currently, the site only ranks for branded terms like "Shofo" and "RecipeBook," which have zero search volume from unacquainted users. To capture commercial intent, the targeting must pivot toward category-level keywords and niche clusters like "hand-object interaction data." Updating the homepage H1 to include "Custom Video Datasets for AI Labs" is a recommended first step.

Evidence
Non-Branded Keyword Rank
0
Keyword Relevance (H1)
Low

29 · Content gaps — Missing 90% of the topical coverage expected for CV data providers

With a critical score of 22/100, the site is effectively invisible to the broader search market. It lacks 90% of the topical coverage expected for a data provider, with zero informational guides or commercial comparison pages. There is no content to intercept users evaluating competitors like Scale AI. The immediate priority is establishing a pillar-and-cluster structure that provides context for the 25M+ clips, moving the site from a "naked" marketplace to an authoritative resource. Creating "Use Case" pages for specific industries like autonomous retail would help bridge this gap.

Evidence
Informational URLs
0
Competitor Comparison Pages
0

30 · Keyword gaps — Competitors dominate 'video dataset' territory while Shofo remains invisible

Shofo is currently conceding all non-branded search volume to incumbents, resulting in a critical score of 18/100. There is zero overlap in shared territory because the site has not targeted category-level keywords. Competitors like Scale AI and Appen maintain a monopoly on "video dataset" queries. Shofo's unique "by the hour" pricing model is its strongest differentiator, but it remains unindexed and unsearched because it lacks dedicated landing pages targeting affordable training data. Launching a targeted campaign around "High-Quality Video Datasets for ML" is required to contest this territory.

Evidence
Shared Keywords with Competitors
0
Category Keyword Presence
None

32 · AI search readiness — Zero JSON-LD schema and missing llms.txt discovery file

AI search readiness is critical at 29/100 as the site is structurally invisible to AI agents. It lacks all forms of structured data, including Organization and Dataset schema, which are required for citations in AI Overviews. The marketplace is a "black box" to crawlers due to its reliance on client-side rendering. Without an llms.txt file or answer-first informational content, AI engines like Perplexity cannot verify the brand's claims or accurately represent its 25M+ clips. Implementing server-side rendering for search results and deploying a machine-readable llms.txt file are the primary recommendations.

Evidence
JSON-LD Schema Blocks
0
AI Accessibility File
404

33 · Fix-priority hygiene — 404 robots.txt and hydration gaps blocking indexation

The site's technical foundation is currently compromised by critical indexation roadblocks that prevent search engines from discovering its 25M+ video clips. The most urgent failure is the robots.txt file, which returns a 404 HTML page containing a "noindex" tag, effectively signaling crawlers to ignore the domain. This is compounded by the total absence of an XML sitemap. Furthermore, a severe "hydration gap" exists within the Next.js architecture; while the raw source contains 557 words of data, the rendered view shows only 64 words because the marketplace grid is hidden behind a client-side "Searching..." skeleton state. To restore visibility, the team must deploy a valid plain-text robots.txt, generate a dynamic sitemap, and implement Server-Side Rendering (SSR) to ensure core content is visible to bots. Additionally, unique, keyword-rich title tags must replace the current generic duplicates across the /report and /contact pages.

Evidence
robots.txt status
404 HTML (noindex)
Rendered word count
64 words
Raw source word count
557 words
Metadata duplication
66%

34 · SEO composite coherence — Technical invisibility concedes 100% of non-branded traffic

RecipeBook by Shofo is a high-potential marketplace currently operating in a state of technical invisibility due to a lack of basic SEO hygiene. While the underlying product—a massive video dataset—is a high-value linkable asset, the site currently concedes all non-branded search traffic to competitors like Scale AI and Labelbox. The coherence of the SEO strategy is undermined by three primary factors: a malformed robots.txt that blocks indexing, a heavy reliance on client-side rendering that creates a "thin content" profile for crawlers, and a total absence of structured data (Schema.org). Without Organization JSON-LD or an llms.txt file, the site is also excluded from AI-driven search results. The 90-day roadmap must prioritize fixing these technical blockers before investing in content. Establishing topical authority will require transitioning from a generic "Shofo" brand signal to targeted keywords like "Computer Vision Training Data" across all metadata and new informational guides.

Evidence
AEO score
29/100
Non-branded visibility
0%
Indexable page count
6

Verdict — 52.9/100: High-value, hidden data

RecipeBook by Shofo is a high-utility marketplace for AI training data that is currently undermined by its technical execution. The site’s strongest asset is its underlying data—25 million video clips—and a clean, modern design stack. However, it fails on basic SEO hygiene, documentation, and trust signals.

Fixable Weaknesses:

  1. Technical indexation (robots.txt and sitemap).
  2. Metadata optimization for high-intent ML keywords.
  3. Public-facing documentation and licensing details.

Who it's for: ML Engineers and Computer Vision Researchers who already know the brand or are referred directly to the marketplace.

Evidence
Overall Score
52.9/100
Authority Score
87/100

90-day roadmap

Window Action Modules Expected effect
Days 1-14 Fix robots.txt 404 and deploy XML sitemap Technical, Health Ensure site discovery and indexation
Days 15-45 Rewrite all Title/H1 tags and implement SSR for the grid On-page, Performance Improve keyword relevance and content depth
Days 46-90 Deploy 'Use Case' pages and Organization Schema Content Gap, AEO Establish topical authority and AI search citations
Evidence
Roadmap Window
90 Days

Methodology & data notes

The review is based on a crawl of 6 pages conducted on 2026-08-04. Several dimensions were excluded: 06, 08, and 18 failed to meet minimum content thresholds, while 12, 13, 22, 28, and 31 were not applicable to this marketplace model. Measurement readiness was assessed via PostHog implementation. For a full explanation of our 34-dimension framework, visit our methodology page.

Evidence
Data Sources
6-page crawl

Questions buyers actually ask

Is RecipeBook by Shofo a reliable source for ML training data?

The site offers a massive repository of 25M+ video clips, but currently lacks the trust signals, such as case studies or enterprise logos, typically expected by professional ML researchers.

Why is RecipeBook not appearing in search results?

Our crawl identified critical technical issues, including a malformed robots.txt file and a lack of unique metadata, which prevent search engines from indexing the marketplace's content.

Does the site support self-service licensing?

No. At the time of review, there is no public-facing documentation or self-serve help center explaining licensing terms or technical specifications for the datasets.

How does RecipeBook compare to competitors like Scale AI?

While the core data asset is significant, RecipeBook lacks the content infrastructure and keyword strategy that established competitors use to dominate the computer vision market.

How this review was made

SiteList examined shofo.ai on August 4, 2026 — pages, screenshots, performance runs, structured data and public records — then scored it across 34 published dimensions. Every claim above cites inspection evidence; nothing is hand-tuned and the verdict is never for sale.

Not assessed in this inspection: 06 · logo-design, 08 · art-direction, 18 · content-brief-authoring. Their weight was redistributed across the assessed dimensions.

Pending enrichment (data we could not fetch this run): SERP lookup for competitor overlap, Similarweb audience overlap, Reddit voice-of-customer sampling, Google People-Also-Ask lookups, Manual screen reader testing of the video search results grid, Keyboard trap verification on the 'Filters' mobile overlay, competitor_crawls, openpagerank

Read the full methodology

53/100RecipeBook by Shofo — Buy video training data by the hour featuring 25M+ clipsJump to review