Services · AI Agent Development
Agents that act,
not just answer.
Goal-directed AI agents that call your tools, reason over context, and route through human approval where wrong actions cost real money. Built on OpenAI, Claude, and the Vercel AI SDK.
What I build
Task agents
Single-purpose agents that complete a named job - schedule a meeting, draft an outreach sequence, triage a ticket, summarize a thread.
Research agents
Agents that explore - web research, document collection, multi-source synthesis with citations.
Operations agents
Agents that run business processes - outreach, categorization, routing, follow-up - with human approval at the high-stakes steps.
Customer-facing assistants
Agents inside your product that perform actions on the user's behalf - schedule, search, generate, integrate. Examples in my shipped work: Caldra AI, DreamCurtains AI, Lindi AI.
The agent stack I default to
- Model layer: Anthropic Claude or OpenAI, with provider failover via the Vercel AI Gateway.
- Orchestration: Vercel AI SDK (TypeScript), or direct SDK use for the control-heavy paths.
- Tools: typed wrappers around your APIs, validated with Zod. MCP server if it fits.
- Memory + RAG: pgvector or Qdrant, hybrid retrieval, citation grounding. See the RAG tutorial.
- HITL: approval queues, confidence-routed escalation, override capture. Pattern catalog in the HITL guide.
- Guardrails: max-step budgets, cost ceilings, output validation, prompt injection defenses.
- Observability: per-step traces, latency, cost per run, evaluation dashboards.
Pricing
| Scope | Timeline | Price |
|---|---|---|
| Single-purpose agent (5-10 tools, eval set, HITL) | 2-5 weeks | $12K-$30K |
| Multi-agent system with shared memory and observability | 4-8 weeks | $25K-$60K |
| Hourly retainer for ongoing iteration | Ongoing | On request |
What this looks like in practice
Worked examples of builds like this: the problem, the workflow, what the AI does and where people stay in the loop.
A second-line support agent that checks the logs before an engineer gets pulled in
Investigates escalated tickets the way an engineer would, from account settings to error logs to known issues, and returns a checked answer or a ready bug report.
Zendesk or Intercom / Jira or Linear / Sentry and Datadog / Admin API or read replica / Slack
A support agent that answers 'where is my order?' from live Shopify and carrier data
Answers order questions on every channel from live Shopify and carrier data, verifies the customer first, and hands refunds and angry customers to a person.
Shopify / Gorgias or Zendesk / WhatsApp Business Platform / Carrier tracking API / Returns app (Loop, ReturnGO)
An agent that chases suppliers for delivery dates and updates the ERP when they change
Emails suppliers about open POs in their language, reads confirmations and date changes, writes agreed dates to the ERP and flags the late parts that will hurt.
Outlook and Microsoft 365 / SAP Business One / NetSuite / Shopify / Microsoft Teams or Slack
Automating portals and legacy systems that have no API, with a person approving the final click
Keys data into insurer portals and API-less desktop systems with a computer-use model, logs every screen, and waits for a person to approve each submission.
Broker management system / Insurer and carrier portals / A legacy desktop ERP / Computer-use model (Copilot Studio, Claude or OpenAI) / Outlook
Freight quote requests in four languages, answered in minutes and priced by your own rules
Reads freight quote requests in your customers' languages, asks for what is missing, prices the lane with your margin rules, and sends or drafts the quote.
Outlook and Microsoft 365 / TMS (Soloplan CarLo, Transporeon, CargoWise or similar) / Freight exchange data (Timocom, Trans.eu; DAT in the US) / HubSpot or the TMS customer record / Truck routing API (PTV or HERE)
An after-hours phone agent for trades that books jobs and puts emergencies through
Answers after-hours calls, puts gas and water emergencies through to the on-call technician, books routine jobs into real slots, and transfers anyone who asks.
Twilio / Real-time voice model / Field service software (HERO, Plancraft, Craftnote; Jobber, ServiceTitan or Housecall Pro in the US) / Google Calendar / WhatsApp
Frequently asked questions
What is an AI agent vs a chatbot?
A chatbot answers. An agent acts. Agents call tools - your APIs, your database, external services - to do real work, in a loop, until a goal is met. They reason about which step to take next instead of returning a single response.
Which agent frameworks do you use?
Whatever fits. Vercel AI SDK for most TypeScript stacks (tool-calling, streaming, generateObject), Anthropic SDK directly for Claude-native agents, OpenAI Assistants for stateful flows, and custom orchestration when frameworks get in the way. I avoid LangChain unless the project benefits from its specific abstractions.
How do you keep agents from going off the rails?
Tight tool schemas, strict output validation, max-step budgets, cost ceilings per run, and human-in-the-loop checkpoints on irreversible actions. See the HITL pattern catalog.
Can agents work with my existing APIs?
Yes - that is the point. I wrap your APIs as typed tools the agent can call. The agent does not need anything special on your side beyond the API existing.
What does AI agent development cost?
Single-purpose agent (one workflow, 5-10 tools, evals): $12K-$30K, 2-5 weeks. Multi-agent system with HITL and observability: $25K-$60K, 4-8 weeks.
Can you ship an agent on top of MCP?
Yes. Model Context Protocol works well when you already have an MCP server (or want one) so tools are reusable across multiple agent runtimes. Happy to wire it up.
See it built: the Caldra AI case study walks through a production agent end to end, and the build guide covers the architecture.
Related: AI integration · AI workflow automation · hire an AI developer in Kosovo