How do I migrate from a direct Responses API call to an Azure AI Foundry Agent?

Camilla Steinsvik 0 Reputation points
2026-10-03T22:56:55.4833333+00:00

Hi,

This is my first software project, and I'm not a professional developer, so please excuse me if I'm misunderstanding how Azure AI Foundry Agents are intended to be used.

I have an Azure AI Foundry Agent that works correctly in Playground.

My application currently uses the Responses API to call a model directly.

The Agent exposes its own Responses Protocol endpoint generated by Azure AI Foundry.

My goal is to have the application use the Agent instead of calling the model directly.

The Agent is already working in Playground and correctly follows the configured system instructions.

My question is:

When an application already uses the Responses API, what is the minimum change required to call a Foundry Agent instead?

  • Do I only need to change the endpoint?
  • Do I need to create an agent session first?
  • Should I use the Agents SDK?
  • Or is there another recommended approach?

Does anyone have a JavaScript/TypeScript example of migrating from direct model calls to an Azure AI Foundry Agent?

Thanks!

Azure API Management
Azure API Management

An Azure service that provides a hybrid, multi-cloud management platform for APIs.

0 comments No comments

3 answers

Sort by: Most helpful
  1. Camilla Steinsvik 0 Reputation points
    2026-10-04T08:02:28.8333333+00:00

    Thank you for taking the time to explain this in such detail. The distinction between the model deployment and the agent endpoint is starting to make much more sense now.

    Looking back at my setup, the model deployment was created first and the application was originally built against that endpoint. The agent was added several weeks later, so there is a real possibility that parts of the application are still pointing to the original model deployment rather than the agent.

    That would actually explain quite a lot. 😅

    I will go through the flow step by step and verify exactly which endpoint is being called from the app. Your explanation has given me a much clearer direction for troubleshooting.

    One additional detail I probably should have mentioned: this is not a chat application.

    The app simply has a button that triggers a single API call to generate a drawing idea. There is no multi-turn conversation, chat history, or user input beyond pressing the button.

    Given that architecture, it sounds even more likely that I should focus on verifying whether that single request is reaching the agent endpoint rather than the original model deployment.

    Thanks again for the time, detail, and references. I really appreciate the help.

    Was this answer helpful?

    0 comments No comments

  2. Senthil kumar 2,580 Reputation points
    2026-10-04T07:14:51.6933333+00:00

    Hi @Camilla Steinsvik

    1. You do not simply replace the model deployment endpoint with the agent endpoint and keep everything else identical. Hosted agents expose their own Responses Protocol endpoint and have additional concepts such as conversations and sessions.
    2. Session creation is not always required.
      • For conversation continuity, the Responses protocol can use either:
      • previous_response_id, or
      • a conversation ID.
      • Reusing a session is primarily for preserving sandbox state, uploaded files, and filesystem contents, not for chat history by itself.
    3. You can call the Agent directly through the Responses Protocol endpoint without adopting the Agents SDK.
      • If your code already uses the OpenAI/Responses API format, the migration effort is typically smaller because the agent endpoint exposes an OpenAI-compatible Responses Protocol surface.
    4. Microsoft's recommended higher-level approach is the Agent Framework/SDK, which handles authentication, tool wiring, and message orchestration for you. However, it is not mandatory if you prefer direct API calls.

    Minimum-change approach

    If your current code looks like this:

    TypeScript

    const response = await client.responses.create({
    model: "gpt-4.1",
    input: "Hello"
    });
    

    Show more lines

    The smallest migration is generally:

    • Change the base URL to the agent's Responses Protocol endpoint.
    • Authenticate as required.
    • Continue calling responses.create(...).
    • Optionally pass a conversation ID or previous_response_id to maintain context across turns. [learn.microsoft.com], [learn.microsoft.com]

    Conceptually:

    TypeScript

    const client = new OpenAI({
    baseURL: AGENT_RESPONSES_ENDPOINT,
    apiKey: process.env.FOUNDRY_API_KEY
    });
     
    const response = await client.responses.create({
    input: "Hello"
    });
    

    In this model, the agent's configured instructions, tools, and behavior come from the Agent definition rather than being supplied in each request. [learn.microsoft.com]

    When do you need sessions?

    Create/manage sessions only if you need capabilities such as:

    • Persistent uploaded files
    • Preserved working directory/state
    • Long-running agent workflows
    • Stateful sandbox reuse across requests

    For normal conversational scenarios, a conversation ID or previous_response_id is usually sufficient.

    Thanks.

    Was this answer helpful?

    0 comments No comments

  3. Salamat Shah 750 Reputation points MVP
    2026-10-04T06:56:47.74+00:00

    Because the Foundry Agent is already created and working in Playground, the application should call the Agent endpoint, rather than continue calling the model deployment directly.

    Do not simply replace the model endpoint URL if the existing code expects a direct model deployment. The request must target the Foundry agent/Responses interface appropriately.

    For a hosted/persisted agent, use the Azure AI Projects SDK. Microsoft recommends the Projects library for working with existing agents, conversations, and responses.

    Create/use a conversation/session when multi-turn conversation history is required. The current TypeScript samples demonstrate reusing an existing agent endpoint, creating a conversation, and generating responses.

    If only an ephemeral agent is required, the application can call the project Responses endpoint at {project_endpoint}/openai/v1/responses; Microsoft identifies this as the recommended Foundry project-scoped Responses API.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.