
Flagship · AE-200
The Answer Engineer
Earns: Certified Answer Engineer
How AI answers are made, and how to get a brand named in them. The core course, and the place to start.
- 12 lessons
- 29 min of video
- Free
Professional course · GEO-300
Earn the Sophyx Certified Technical GEO Engineer – Professional credential.
Technical GEO Engineer is a free 12-lesson video course by Sophyx for developers and technical SEOs. You learn to build a site that AI systems can fetch, read, trust and act on, and to keep it that way: verified bot access, facts in the first HTML response, structured data, pages and APIs that AI agents can use, and checks on every deploy. It builds on The Answer Engineer, and every lesson works through one fictional company: Gearwick Bikes, an online bike shop.
You own the code and the deploys. You want AI fetchers to get the facts in the first response, and checks that stop a release from breaking it.
You know crawling and indexing. You want to know which AI bots visit, how to check they are real, and how to set a bot policy for each job.
You run the servers and the CDN. You want verified bot logs, alerts when AI bots get errors, and responses that stay fast.
The Answer Engineer (AE-200) first is recommended, plus comfort with HTTP, HTML and a terminal. If you already know how AI search works, you can start here.
The lessons follow Gearwick Bikes, a fictional online bike shop whose storefront is a JavaScript app, and its fictional rival Tarnway Cycles. Every brand in the course is fictional and every number is illustrative.
A sample of the 36 objectives across the 4 modules.
Twelve lessons in four modules: let AI in, make the site machine-readable, get it ready for agents, then observe and operate.
4 modules, 12 lessons, about 36 minutes of video. Each lesson has captions, chapters, a transcript and a practice quiz.
Module 1 · 3 lessons ·
The three jobs AI bots do and how to verify a bot before you trust it; bot policy as code (robots.txt by job, control tokens, CDN rules, proposals labelled); what a fetcher without JavaScript actually receives.

Lesson 1
AI companies run bots for three jobs: training crawlers, search indexers, and user-triggered fetchers that fetch a page the moment a person asks an assistant something. Each bot names itself with a product token inside its user agent string; Google-Extended and Applebot-Extended are control tokens, not crawlers. A token is only text, so verify a bot before you trust it: the vendor's published IP list, reverse then forward DNS, or a signed request (Web Bot Auth, an IETF draft). Bot names and behaviours are as of October 2026.

Lesson 2
Bot policy is one decision per job, written as code and tested. robots.txt (RFC 9309) groups rules by product token, applies the longest matching path, and can announce your sitemap; most vendors cache it for about a day. Decide per job: search indexers and user-triggered fetchers bring answers and visits; training is a business choice; control tokens such as Google-Extended set AI use without touching the crawl. Your CDN or WAF adds a second layer: allow verified bots by category, block the unverified. Newer signals (Content Signals, the IETF AIPREF drafts) are proposals, and some user fetchers say robots.txt may not apply to them. Keep the policy in version control with a test for every allowed token.

Lesson 3
A fetcher receives one HTTP response. Googlebot and Applebot render JavaScript; the crawlers and user fetchers from OpenAI, Anthropic and Perplexity did not in the largest published study (Vercel and MERJ, 2024), and none of them documents rendering as of October 2026, so a page whose facts arrive only through scripts is empty to them. (Browser agents are a different case: Lesson 07.) Test it: fetch the page as the bot with JavaScript off and read the HTML. Serve the facts in that first response with server-side rendering, static generation or pre-rendering, and keep the status honest: a 200 for real pages, a real 404 for missing ones, no consent walls, interstitials or geo-blocks in front of bots, and no "soft 404" shells.
Module 2 · 3 lessons ·
Structured data as an entity graph with stable @ids; HTML written as quotable passages; discovery files and feeds judged by what they earn.

Lesson 4
Structured data is a graph, not a sprinkle of snippets. In JSON-LD, give every real-world thing one node with a stable @id, reuse that @id on every page that mentions it, and connect the nodes: the organisation, its website, its locations, its products and offers, its people. Link each entity to its official profiles with sameAs. Keep the markup true to the visible page, because Google ignores or penalises markup for things users cannot see, and validate it on every deploy. Google says no new markup is required for its AI features (as of 2026); the graph still removes ambiguity about who you are.

Lesson 5
AI systems quote passages, so write the page as passages. One question per section, answered in the first sentence under a heading that says the question; article and section elements with a clean h1–h6 order; tables for specifications and comparisons; ordered lists for steps; time elements for dates; a lang attribute; link text that names its target. Text inside tabs and accordions counts when it is in the HTML, not when it is loaded on click. Specs in images, PDFs and carousels are invisible: put them in text.

Lesson 6
Discovery files tell bots what you have and what changed. A sitemap lists canonical URLs with an honest lastmod (50,000 URLs and 50 MB per file; index files above that), announced with a Sitemap line in robots.txt. IndexNow pushes changed URLs to Bing, Yandex, Naver, Seznam, Yep and Amazon, but not to Google. Feeds (RSS, Atom) and HTTP validators (ETag, Last-Modified) tell fetchers what is fresh. llms.txt is a proposal: a markdown map of your key pages; as of 2026 no major assistant documents reading it when it answers, and Google says its AI features need no AI text files, so publish it cheaply and judge it by the fetches in your own logs.
Module 3 · 3 lessons ·
Pages agents can operate; one source of truth exposed as an API and an MCP server; fast, cacheable responses that never shrink the crawl.

Lesson 7
AI agents (as of October 2026: ChatGPT's cloud browser, Claude's computer use, Google-Agent) act through the page the way assistive technology does: they read the accessibility tree and operate real controls. Give them real forms with labels, buttons that are buttons, links with names, visible state and error messages, keyboard paths with no hover-only menus, and a checkout that does not hide steps behind scripts or unlabelled icons. Let verified agents through your bot defences, and keep consent walls off the task path. The test is the walkthrough: find, compare, act.

Lesson 8
Assistants and agents can call you directly. Expose one source of truth as an API: an OpenAPI description, read-only endpoints for facts such as price, stock, hours and specs, and the same data your pages show. The Model Context Protocol (MCP) standardises how an assistant connects: a server offers tools and resources over Streamable HTTP, with OAuth 2.1 for anything personal, and consent before any action. Rate-limit, log and version it, and keep it in sync with the site, because two sources of truth make two different answers. App directories are built on MCP servers (as of October 2026, ChatGPT's plugin directory, which replaced its app directory in July 2026); treat their rules as dated.

Lesson 9
Fetchers give up fast. Keep time to first byte low, serve bots from the edge cache, and use Cache-Control, ETag and 304 so unchanged pages cost nothing. Google shrinks its crawl when responses slow down or return 5xx and 429, and a robots.txt that errors can hold crawling back altogether. Rate limits should spare verified bots; redirects should land in one hop; Core Web Vitals (LCP 2.5 s, INP 200 ms, CLS 0.1 at the 75th percentile) keep the human side fast. Monitor with synthetic fetches as the bots.
Module 4 · 3 lessons ·
A verified bot log pipeline with alerts; referral plumbing that keeps the AI source attached; GEO checks in CI with owners and a change log.

Lesson 10
Turn raw logs into a verified bot feed. Parse CDN or server logs, classify each request by product token and job, then verify: match the IP against the vendor lists you refresh daily, do reverse and forward DNS with a cache, and check signatures where bots sign. Store the result with the status code and the path. Then alert on what matters to an engineer: 403, 429 or 5xx sent to a verified bot, a new token appearing, a key page no longer fetched, robots.txt fetches failing. Keep IP retention short. (GEO-310 reads the same feed for metrics; this lesson builds it.)

Lesson 11
A click from an AI answer only shows as an AI visit if the plumbing keeps the source attached. Browsers default to strict-origin-when-cross-origin, sending only the origin across sites and nothing on an HTTPS-to-HTTP step, so serve HTTPS everywhere. ChatGPT appends utm_source=chatgpt.com; a redirect that drops the query string, a canonicalisation step, a consent wall or a geo-redirect to the home page loses it. Make sure landing pages exist, log the referrer and the UTM on the server, and test the path from an answer link to the final page. GA4's AI Assistant channel (as of 2026) can only see what survives.

Lesson 12
Everything in this course can be undone by one deploy. Put the checks in CI: robots.txt allow tests per token, a fetch of each key page as a bot with JavaScript off that fails on an empty shell, structured-data validation against the entity graph, sitemap validity and lastmod sanity, a TTFB budget, a referral test, and a check that staging rules never ship. Give each policy an owner, keep a change log, and review the vendor documentation quarterly, because tokens and rules change. GEO is operated, not launched.
Watch, practice, prove, then get certified: the same four steps in every Xlearn by Sophyx course.
Watch all 12 lessons. Each one is about 2 to 3 minutes, with captions, chapters and a transcript.
Take the practice quiz after each lesson. Try as often as you like, and every wrong answer tells you why.
Pass the final exam: 40 questions, 60 minutes, 75% to pass, and 24 hours between attempts. Then a final project, scored against a written rubric with feedback. You can revise and resubmit.
Your certificate gets its own credential ID and a public page anyone can check. Add it to LinkedIn in one click or download the PDF. It is valid for 12 months.
You can retake it after 24 hours.
Pick a brand: your own, a client's with their permission, or the fictional Gearwick Bikes data pack. Hand in one deliverable for each part below. Each criterion is scored from 0 to 3 against its descriptor, and you pass with a 2 or higher on every one. Two graders score one in ten submissions, to keep scoring fair. You get written feedback and can revise and resubmit.
Verified bot inventory (M1)
One week of logs: AI requests grouped by product token and job, each top token verified by IP list, reverse DNS or signature.
Policy as code (M1)
A decision table by job, expressed as robots.txt and a CDN or WAF rule, in version control, with a test that fetches robots.txt as each allowed token.
Raw-HTML audit (M1)
The 10 key pages fetched as a bot with JavaScript off; every missing fact listed; fixes that put the facts in the first response.
Entity graph (M2)
JSON-LD with one node per entity, stable @ids reused across pages, sameAs links, validated, and true to the visible pages.
Passage pages (M2)
Five pages restructured as one-answer sections: question headings, direct first sentences, tables for specs, lists for steps, clean heading order.
Discovery set (M2)
A sitemap with honest lastmod and an index if needed, the Sitemap line in robots.txt, IndexNow or a feed, and llms.txt labelled as a proposal with its fetches measured.
Agent walkthrough (M3)
An agent given one task (find, compare, act), each failure logged with its cause, the three worst controls fixed.
Facts API or MCP (M3)
An OpenAPI description or an MCP tool and resource schema for the ten facts assistants ask about, with auth, rate limits and a sync test against the pages.
Performance and caching plan (M3)
TTFB measured for ten pages as a bot, ETag or Last-Modified with confirmed 304s, every 5xx and 429 sent to verified bots found and addressed.
Verified bot feed with alerts (M4)
A pipeline that parses, classifies and verifies bot requests daily, a dashboard of verdicts and status codes, and two alerts with owners.
Referral plumbing test (M4)
Five AI answer links traced to the landing page; referrer and UTM checked; every redirect or wall that drops them fixed.
GEO CI pipeline (M4)
Five checks in CI (robots, raw HTML, JSON-LD, sitemap, TTFB) that block a failing deploy, a named owner and a dated change log.
Honesty and labelling
Fictional brands and illustrative numbers labelled; platform specifics dated; proposals named as proposals; limits stated.
One page that proves what you learned, with a link anyone can open to check it.

Xlearn by Sophyx
Certificate of completion
This certifies that
Your name
has earned the credential
Sophyx Certified Technical GEO Engineer – Professional
for completing the 12 lessons of Technical GEO Engineer, passing the final exam and passing the final project.
It is a Sophyx credential, not an accredited qualification.
Open a sample certificateMore on how certificates work is on the Xlearn by Sophyx page.
Every course sits on one ladder, from fundamentals to expert. See the courses coming next.

Flagship · AE-200
Earns: Certified Answer Engineer
How AI answers are made, and how to get a brand named in them. The core course, and the place to start.

Associate · AEO-210
Earns: Sophyx Certified AEO Content Specialist – Associate
For content teams: plan, write, prove, refresh and measure pages that AI assistants quote and send buyers to.

Professional · GEO-310
Earns: Sophyx Certified GEO Analytics Professional
For analytics and growth teams: measure AI visibility from crawler reads to conversions, with numbers you can defend.
Twelve short lessons, free. Finish with a certificate anyone can check.