Most companies shopping for an "AI implementation partner" have never bought this kind of work before, and the market is flooded with agencies that rebranded overnight from generic software shops to "AI consultancies." The questions below won't guarantee a good outcome, but they'll filter out firms that can't actually do the work.

Start with the use case, not the vendor

Before you talk to anyone, write down the specific business problem you're trying to solve and what "done" looks like in numbers you already track — cycle time, error rate, cost per ticket, conversion rate. A firm that jumps straight to "we can build you an AI agent" without asking what success looks like is selling a technology, not solving a problem. The best partners push back on vague requests and help you scope down to a single measurable pilot before discussing a bigger engagement.

Questions to ask in the first call

  • Show me something you shipped, not a deck. Ask for a live product, a demo environment, or a case study with real technical detail — which models, what data pipeline, what evaluation approach. Case studies with no technical specifics are a red flag.
  • Who exactly will work on my project? Get names and backgrounds, not just the sales team. Many firms pitch senior architects and staff junior generalists on delivery.
  • What happens when the model is wrong? A serious partner should have a concrete answer involving evaluation harnesses, guardrails, human-in-the-loop review, and monitoring — not just "we'll fine-tune it."
  • How do you handle my data? Ask where it goes, whether it trains any shared or third-party models, and what your break clause looks like if you're not comfortable with the answer.
  • What's the smallest version of this you'd build first? A partner who can't propose a scoped-down pilot either hasn't done this enough times, or is optimizing for a bigger invoice.

Red flags

  • Guaranteed ROI numbers or fixed percentages promised before they've seen your data or workflows.
  • Reluctance to name the actual foundation models, frameworks, or infrastructure they use.
  • No mention of evaluation, testing, or monitoring — only "deployment."
  • A sales process that skips discovery and jumps straight to a statement of work.
  • Case studies that cite only client logos with no description of what was actually built.
  • Team bios that are entirely business/strategy background, with no one who can speak to model architecture, data pipelines, or infra at a technical level.

Scoping a pilot correctly

A good first engagement is time-boxed (typically four to ten weeks), tied to one workflow or use case, and has an explicit exit ramp — you should be able to walk away after the pilot with working software or a clear no-go decision, not a half-built system that only makes sense if you keep paying the same firm indefinitely. Ask upfront what you own if you end the engagement after the pilot: code, models, prompts, evaluation data, and documentation should all transfer to you, not stay locked in the vendor's tooling.

Team composition that actually ships

A team that can implement AI in production typically covers four roles: someone who owns data engineering (getting your data into a usable, permissioned state is usually the bulk of the real work, not the model itself); an ML/AI engineer who can weigh build-vs-buy rather than just wiring up an API call; someone accountable for evaluation and monitoring after launch, since AI systems drift and "we'll hand it off, you're on your own" is a common, costly gap; and a product or domain owner who knows when the AI's output is actually useful versus technically correct but useless.

Boutique shops and large consultancies both have a place here. A boutique implementation shop — firms like asaasin.ai, which focuses on embedded AI engineering teams for product companies, are one example — can move faster and put senior people directly on your project, since there are fewer layers between you and the work. A larger firm may make more sense when you need broad geographic coverage, compliance depth, or the ability to staff a much bigger team quickly. Neither size is inherently better; match the firm's shape to your problem's shape.

IP and data ownership — get this in writing

Before signing, confirm in the contract itself, not just a sales conversation: you own the code, prompts, fine-tuned weights, and evaluation datasets produced during the engagement; your data is never used to train models shared with other clients unless you've explicitly opted in; there's a clear data retention and deletion policy once the engagement ends; and if the vendor relies on proprietary internal tooling, you understand what breaks if you stop paying them.

Judging technical depth without being technical yourself

Ask the team to walk you through a past project's failure mode — what didn't work, how they detected it, what they changed. Teams with real production experience have a specific, unglamorous story here (a RAG pipeline that hallucinated on edge cases, a model that drifted after a data schema change). Teams without real depth will dodge the question or describe only successes. Then get a second opinion from someone technical who isn't invested in the deal — a friendly CTO, an advisor, even a competing vendor's free discovery call. The best filter against overselling is someone who can ask the sharp follow-up question you wouldn't know to ask yourself.