ai-rag
ai-rag 插件为大语言模型提供了检索增强生成(RAG)能力。它支持从外部数据源高效检索相关文档或信息,用于增强大语言模型的响应,从而提高生成输出的准确性和上下文相关性。
该插件支持使用 Azure OpenAI 和 Azure AI Search 服务来生成嵌入和执行向量搜索。
示例
要跟随示例进行操作,请先创建一个 Azure 账户 并完成以下步骤:
- 在 Azure AI Foundry 中部署一个生成式对话模型(如
gpt-4o)和一个嵌入模型(如text-embedding-3-large)。获取 API Key 和模型端点。 - 按照 Azure 的示例 使用 Python 在 Azure AI Search 中准备向量搜索。该示例将创建一个名为
vectest的搜索索引,并包含所需的架构,同时上传包含 108 条各类 Azure 服务描述的示例数据,以便根据title和content生成嵌入titleVector和contentVector。在 Python 中执行向量搜索之前,请完成所有设置。 - 在 Azure AI Search 中,获取 Azure 向量搜索 API Key 和搜索服务端点。
将 API Key 和端点保存到环境变量中:
# 替换为你的配置值
AZ_OPENAI_DOMAIN=https://ai-plugin-developer.openai.azure.com
AZ_OPENAI_API_KEY=YOUR_AZURE_OPENAI_API_KEY
AZ_CHAT_ENDPOINT=${AZ_OPENAI_DOMAIN}/openai/deployments/gpt-4o/chat/completions?api-version=2024-02-15-preview
AZ_EMBEDDING_MODEL=text-embedding-3-large
AZ_EMBEDDINGS_ENDPOINT=${AZ_OPENAI_DOMAIN}/openai/deployments/${AZ_EMBEDDING_MODEL}/embeddings?api-version=2023-05-15
AZ_AI_SEARCH_SVC_DOMAIN=https://ai-plugin-developer.search.windows.net
AZ_AI_SEARCH_KEY=YOUR_AZURE_AI_SEARCH_API_KEY
AZ_AI_SEARCH_INDEX=vectest
AZ_AI_SEARCH_ENDPOINT=${AZ_AI_SEARCH_SVC_DOMAIN}/indexes/${AZ_AI_SEARCH_INDEX}/docs/search?api-version=2024-07-01
集成 Azure 以生成 RAG 增强响应
以下示例演示了如何使用 ai-proxy 插件代理请求到 Azure OpenAI 大语言模型,并使用 ai-rag 插件生成嵌入并执行向量搜索,以增强大语言模型的响应。
- Admin API
- ADC
- Ingress Controller
创建路由如下:
curl "http://127.0.0.1:9180/apisix/admin/routes" -X PUT \
-H "X-API-KEY: ${ADMIN_API_KEY}" \
-d '{
"id": "ai-rag-route",
"uri": "/rag",
"plugins": {
"ai-rag": {
"embeddings_provider": {
"azure_openai": {
"endpoint": "'"$AZ_EMBEDDINGS_ENDPOINT"'",
"api_key": "'"$AZ_OPENAI_API_KEY"'"
}
},
"vector_search_provider": {
"azure_ai_search": {
"endpoint": "'"$AZ_AI_SEARCH_ENDPOINT"'",
"api_key": "'"$AZ_AI_SEARCH_KEY"'"
}
}
},
"ai-proxy": {
"provider": "openai",
"auth": {
"header": {
"api-key": "'"$AZ_OPENAI_API_KEY"'"
}
},
"model": "gpt-4o",
"override": {
"endpoint": "'"$AZ_CHAT_ENDPOINT"'"
}
}
}
}'
创建一个配置了 ai-rag 和 ai-proxy 插件的路由,如下所示:
adc.yaml
services:
- name: ai-rag-service
routes:
- name: ai-rag-route
uris:
- /rag
methods:
- POST
plugins:
ai-rag:
embeddings_provider:
azure_openai:
endpoint: "${AZ_EMBEDDINGS_ENDPOINT}"
api_key: "${AZ_OPENAI_API_KEY}"
vector_search_provider:
azure_ai_search:
endpoint: "${AZ_AI_SEARCH_ENDPOINT}"
api_key: "${AZ_AI_SEARCH_KEY}"
ai-proxy:
provider: openai
auth:
header:
api-key: "${AZ_OPENAI_API_KEY}"
model: gpt-4o
override:
endpoint: "${AZ_CHAT_ENDPOINT}"
将配置同步到网关:
adc sync -f adc.yaml
- Gateway API
- APISIX CRD
创建一个配置了 ai-rag 和 ai-proxy 插件的路由,如下所示:
ai-rag-ic.yaml
apiVersion: apisix.apache.org/v1alpha1
kind: PluginConfig
metadata:
namespace: aic
name: ai-rag-plugin-config
spec:
plugins:
- name: ai-rag
config:
embeddings_provider:
azure_openai:
endpoint: "https://ai-plugin-developer.openai.azure.com/openai/deployments/text-embedding-3-large/embeddings?api-version=2023-05-15"
api_key: "YOUR_AZURE_OPENAI_API_KEY"
vector_search_provider:
azure_ai_search:
endpoint: "https://ai-plugin-developer.search.windows.net/indexes/vectest/docs/search?api-version=2024-07-01"
api_key: "YOUR_AZURE_AI_SEARCH_API_KEY"
- name: ai-proxy
config:
provider: openai
auth:
header:
api-key: "YOUR_AZURE_OPENAI_API_KEY"
model: gpt-4o
override:
endpoint: "https://ai-plugin-developer.openai.azure.com/openai/deployments/gpt-4o/chat/completions?api-version=2024-02-15-preview"
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
namespace: aic
name: ai-rag-route
spec:
parentRefs:
- name: apisix
rules:
- matches:
- path:
type: Exact
value: /rag
method: POST
filters:
- type: ExtensionRef
extensionRef:
group: apisix.apache.org
kind: PluginConfig
name: ai-rag-plugin-config
将配置应用到集群:
kubectl apply -f ai-rag-ic.yaml
创建一个配置了 ai-rag 和 ai-proxy 插件的路由,如下所示:
ai-rag-ic.yaml
apiVersion: apisix.apache.org/v2
kind: ApisixRoute
metadata:
namespace: aic
name: ai-rag-route
spec:
ingressClassName: apisix
http:
- name: ai-rag-route
match:
paths:
- /rag
methods:
- POST
plugins:
- name: ai-rag
enable: true
config:
embeddings_provider:
azure_openai:
endpoint: "https://ai-plugin-developer.openai.azure.com/openai/deployments/text-embedding-3-large/embeddings?api-version=2023-05-15"
api_key: "YOUR_AZURE_OPENAI_API_KEY"
vector_search_provider:
azure_ai_search:
endpoint: "https://ai-plugin-developer.search.windows.net/indexes/vectest/docs/search?api-version=2024-07-01"
api_key: "YOUR_AZURE_AI_SEARCH_API_KEY"
- name: ai-proxy
enable: true
config:
provider: openai
auth:
header:
api-key: "YOUR_AZURE_OPENAI_API_KEY"
model: gpt-4o
override:
endpoint: "https://ai-plugin-developer.openai.azure.com/openai/deployments/gpt-4o/chat/completions?api-version=2024-02-15-preview"
将配置应用到集群:
kubectl apply -f ai-rag-ic.yaml
向该路由发送一个 POST 请求,请求体中包含向量字段名称、嵌入模型维度和输入提示词:
curl "http://127.0.0.1:9080/rag" -X POST \
-H "Content-Type: application/json" \
-d '{
"ai_rag":{
"vector_search":{
"fields":"contentVector"
},
"embeddings":{
"input":"Which Azure services are good for DevOps?",
"dimensions":1024
}
}
}'
你应该收到类似于以下的 HTTP/1.1 200 OK 响应:
{
"choices": [
{
"content_filter_results": {
...
},
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "Here is a list of Azure services categorized along with a brief description of each based on the provided JSON data:\n\n### Developer Tools\n- **Azure DevOps**: A suite of services that help you plan, build, and deploy applications, including Azure Boards, Azure Repos, Azure Pipelines, Azure Test Plans, and Azure Artifacts.\n- **Azure DevTest Labs**: A fully managed service to create, manage, and share development and test environments in Azure, supporting custom templates, cost management, and integration with Azure DevOps.\n\n### Containers\n- **Azure Kubernetes Service (AKS)**: A managed container orchestration service based on Kubernetes, simplifying deployment and management of containerized applications with features like automatic upgrades and scaling.\n- **Azure Container Instances**: A serverless container runtime to run and scale containerized applications without managing the underlying infrastructure.\n- **Azure Container Registry**: A fully managed Docker registry service to store and manage container images and artifacts.\n\n### Web\n- **Azure App Service**: A fully managed platform for building, deploying, and scaling web apps, mobile app backends, and RESTful APIs with support for multiple programming languages.\n- **Azure SignalR Service**: A fully managed real-time messaging service to build and scale real-time web applications.\n- **Azure Static Web Apps**: A serverless hosting service for modern web applications using static front-end technologies and serverless APIs.\n\n### Compute\n- **Azure Virtual Machines**: Infrastructure-as-a-Service (IaaS) offering for deploying and managing virtual machines in the cloud.\n- **Azure Functions**: A serverless compute service to run event-driven code without managing infrastructure.\n- **Azure Batch**: A job scheduling service to run large-scale parallel and high-performance computing (HPC) applications.\n- **Azure Service Fabric**: A platform to build, deploy, and manage scalable and reliable microservices and container-based applications.\n- **Azure Quantum**: A quantum computing service to build and run quantum applications.\n- **Azure Stack Edge**: A managed edge computing appliance to run Azure services and AI workloads on-premises or at the edge.\n\n### Security\n- **Azure Bastion**: A fully managed service providing secure and scalable remote access to virtual machines.\n- **Azure Security Center**: A unified security management service to protect workloads across Azure and on-premises infrastructure.\n- **Azure DDoS Protection**: A cloud-based service to protect applications and resources from distributed denial-of-service (DDoS) attacks.\n\n### Databases\n",
"role": "assistant"
}
}
],
"created": 1740625850,
"id": "chatcmpl-B54gQdumpfioMPIybFnirr6rq9ZZS",
"model": "gpt-4o-2024-05-13",
"object": "chat.completion",
"prompt_filter_results": [
{
"prompt_index": 0,
"content_filter_results": {
...
}
}
],
"system_fingerprint": "fp_65792305e4",
"usage": {
...
}
}