Skip to content
Proudly based in Nova Scotia, Canada · clients welcome from every countryContact usClient login
CodeLumaDevelopment Inc.

Home / Blog / Article

CodeLuma insights · September 21, 2026 · 9 min read

Is Your Website Ready for Modern Search? A Checklist

More people ask an assistant before opening a search page. A practical checklist: crawler access, server-rendered text, structured data, llms.txt and sitemaps.

More and more people now ask an AI assistant a question before they ever open a search results page: "Who builds custom software in Nova Scotia?", "Which courier covers Wolfville to Halifax?", "What is a good free memory game for my kid?" If your website is not something those systems can read, understand and trust, you are simply not part of the answer. The good news is that most of what helps is not exotic. It is clear writing, honest structure and a few files that tell machines what your site is. This is the checklist we use, and what we found when we ran it against our own sites.

What "making sense to AI" actually means

AI search tools and assistants do three things with a website: they crawl it (fetch the pages), they understand it (work out what the business or product is, what each page says and how the facts fit together) and they decide whether to use it (quote it, link to it or recommend it). A site can fail at any of the three. It might block the crawler by accident, hide its content behind scripts that never run, or say the same thing three different ways on three different pages so that nothing reads as a clear fact.

One honest caveat before we start: nobody can promise that an AI assistant will cite your site. Anyone who guarantees that is guessing. What you can do is remove every avoidable reason you would be skipped, and make your real expertise easy to read. That is the whole job.

A processor chip on a circuit board, close up
Original CodeLuma 3D render: a processor chip, standing in for the systems that read your website.

Step 1: let the right crawlers in

Start with the boring part, because it is where the most damage happens. Your robots.txt file tells crawlers what they may fetch. Many sites block everything unfamiliar after a security plugin or a hosting default did it for them, and never notice. Check that the crawlers that power search and assistants are allowed, and that only genuinely private areas (admin screens, private uploads, checkout and account pages) are off limits.

The crawlers worth knowing about include OpenAI's GPTBot, OAI-SearchBot and ChatGPT-User, Anthropic's ClaudeBot and Claude-SearchBot, PerplexityBot, Google-Extended, Applebot-Extended and Bingbot. Allowing them is a business decision, not a technical one: allowing them means your content can be read and referenced. Blocking them is a legitimate choice for some publishers, but it should be a choice, not an accident.

User-agent: GPTBot
Allow: /
Disallow: /admin/
Disallow: /api/

User-agent: ClaudeBot
Allow: /
Disallow: /admin/
Disallow: /api/

Sitemap: https://example.com/sitemap.xml

Also confirm that the pages themselves are reachable: no login walls on public content, no "noindex" left over from a staging site, and no firewall rule that returns errors to unfamiliar visitors. A quick test is to fetch a page with a plain command-line tool and read what comes back.

Step 2: put the facts in plain, server-rendered text

If your important content only appears after a script runs in the browser, many crawlers will see an empty shell. Make sure the essential facts are in the HTML the server sends: who you are, what you do, where you do it, what you offer, and how to get in touch. Use one clear heading per page, descriptive subheadings, short paragraphs and real sentences. "We build custom software, websites and Android and iPhone apps from Nova Scotia for clients across Canada and around the world" is something a machine can quote. A slogan and a stock photo are not.

Write the way a knowledgeable person would explain the page out loud. Answer the obvious questions on the page itself: what is it, who is it for, what does it cost or how is it priced, how long does it take, what happens next. A visible FAQ section is one of the most useful things you can add, for people and for machines.

Step 3: add structured data that matches the page

Structured data (usually JSON-LD using the schema.org vocabulary) is a labelled summary of a page that machines can read without guessing. The types that matter for most small and medium businesses are:

  • Organization or LocalBusiness: name, address, phone, service area, logo and links to your official profiles.
  • WebSite and WebPage: what the site is and what each page is about.
  • FAQPage: the questions and answers that are visibly on the page.
  • Service, Product or SoftwareApplication: what you sell, with honest descriptions.
  • BreadcrumbList: how the page fits into the site.
  • Article or BlogPosting: for guides like this one, with author and date.

The rule that keeps you out of trouble: structured data must describe what is actually on the page. Do not mark up reviews you do not have, prices you do not charge or an FAQ nobody can see. Search engines penalize that, and it undermines trust in everything else you publish.

A network patch panel with colourful cables plugged into glowing ports
Original CodeLuma 3D render: a network patch panel, a picture of connected things that need to agree with each other.

Step 4: add an llms.txt file, and be realistic about it

llms.txt is a proposed convention: a plain text file at the root of your site (for example yoursite.com/llms.txt) that gives AI assistants a short, human-readable guide to the site. A good one has a one-paragraph summary of who you are, a list of your most important pages with a one-line description each, and the key facts (location, services, currency, how to contact you).

Be realistic: it is not an official ranking factor, and not every assistant reads it today. But it costs almost nothing, it forces you to write down your own key facts clearly, and it gives any tool that does look for it a clean, current summary instead of a guess. The one real risk is letting it go stale. An out-of-date llms.txt that describes services you no longer offer is worse than none. If your site changes often, generate the file from the same data as the site so it can never drift.

Step 5: make the facts agree everywhere

Machines build a picture of your business by comparing what different sources say. If your website says "Halifax", your Google Business Profile says "Dartmouth", a directory lists an old phone number and your social pages use a slightly different company name, the picture gets blurry. Pick one exact business name, one address format, one phone number and one description, and use them everywhere: site, structured data, llms.txt, business profiles, social accounts and directories. Link your official profiles from your Organization markup so it is obvious they belong together.

Step 6: keep it fast, fresh and easy to discover

  • Sitemaps: an accurate XML sitemap with real last-modified dates, listed in robots.txt and submitted in Google Search Console and Bing Webmaster Tools.
  • IndexNow: a simple way to tell participating search engines the moment a page is added or changed, instead of waiting to be recrawled.
  • Speed: slow pages get fetched less and abandoned more. Compress images, cache properly and avoid heavy scripts on content pages.
  • Freshness: publish and update real content. A guide with a current date and specific detail beats a thin page written to hit a keyword.
  • Canonical links and clean redirects: one address per page, and permanent (301) redirects when something moves so old links and earned trust carry over.

What we found when we audited our own sites

We build and look after several sites of our own, so we ran this checklist across six of them: CodeLuma, Nerdania, Smart Marty Sky, Hants County Express, HamRadioList and WalkieTalkieHRL. The results were instructive, and a little humbling:

  • None of the six blocked AI crawlers. That was reassuring, and it is the most important check.
  • Two sites had no llms.txt at all, and a third had one that no longer described our work accurately. Our own portfolio had moved on and the file had not.
  • One site had only basic structured data and no Organization or FAQ markup, even though the page already had plenty of useful information to describe.
  • Names had drifted: a product we had renamed still appeared under its old name in a couple of places, including a payments profile.

The fixes were small. For our free games site, Nerdania, we made the llms.txt file generate itself from the live game catalogue, so it lists every game and category and updates whenever a game is added. We added an Organization and FAQ block to its home page with a matching visible FAQ, made the crawler rules explicit, and updated our own llms.txt to reflect the current portfolio. None of it took long, and none of it changes how the site looks. That is typical: the work is mostly clarity, not redesign.

Common mistakes to avoid

  • Blocking all bots because a security tool suggested it, without deciding what you actually want to allow.
  • Marking up content that is not on the page, such as invented reviews, ratings or FAQs.
  • Relying on a script-only page for critical content.
  • Stuffing the same keyword everywhere instead of answering the real question.
  • Letting llms.txt, sitemaps and structured data go stale after a redesign.
  • Assuming a plugin "handles AI SEO". Plugins can help, but the facts still have to be yours and still have to be true.

The checklist

  1. robots.txt allows the crawlers you want, and only private areas are blocked.
  2. Public pages load without a login and are not marked noindex.
  3. Key facts appear in the server-sent HTML: who, what, where, how to contact.
  4. Each page has one clear heading, a descriptive title and a real meta description.
  5. Structured data (Organization or LocalBusiness, WebSite, FAQPage, Service or Product) matches visible content.
  6. A visible FAQ answers the questions customers actually ask.
  7. llms.txt exists, is current and links to your most important pages.
  8. Name, address, phone and description agree across the site, profiles and directories.
  9. Sitemap is accurate, listed in robots.txt and submitted; IndexNow is set up.
  10. Pages are fast, canonical and free of broken links and chained redirects.
  11. You review all of the above whenever the site changes.

For Nova Scotia businesses

Local businesses have an advantage here, because the questions people ask are specific: "same-day courier Wolfville to Halifax", "web developer near Truro", "plumber in the Annapolis Valley". Make sure every service page names the towns and regions you actually serve, that your Google Business Profile matches your website exactly, and that your structured data lists your service area. Assistants answering a local question look for a business whose facts are clear, consistent and current. Be that business.

Need this done properly?

CodeLuma Development Inc. is a full-stack studio in Nova Scotia. We build websites and software that are fast, secure and readable to both people and machines, and we can audit an existing site against this checklist and fix what we find. You can see the kind of work we do in our portfolio, read about our web development services, or start a project and tell us what you are building. If it is a complex, connected system, or you intend your business to grow into something significant, that is exactly the kind of project we like.

Share this article: Facebook · LinkedIn · X · Email

← All articles

Ready to put this into practice?

Talk to a Nova Scotia full-stack team that builds complex, connected systems for clients across Canada and worldwide.

Start a project