LLM Usage Insights

In LLM Usage Dashboard, you can monitor token usage on LLMs.

Prerequisites

  • At least one active search client.

  • At least one active LLM integration with OpenAI

Access LLM Usage Dashboard

  1. Expand Large Language Model and open Usage Dashboard.

  2. Use the Search Client and Date Range filters to generate a report.

Overview: Total Requests

The Total Requests report contains the following five tiles:

Fig. A snapshot of the Total Requests report in LLM Usage Insights.

  • Input Tokens counts the total number of tokens sent from SearchUnify to an LLM for the selected search client and date range. When multiple LLMs are connected, the tokens sent to each LLM are counted separately at the bottom of the tile.

  • Generated Tokens counts the total number of tokens an LLM uses to generate all responses for the selected search client and date range. When multiple LLMs are connected, the tokens used by each LLM are counted separately at the bottom of the tile.

  • Average LLM Response Time is the total time spent generating responses divided by the total number of requests.

  • Error Rate is the ratio of requests that resulted in errors divided by the total number of requests.

  • Integration is visible only when multiple LLM integrations are in use.

API Consumption

Each SearchUnify customer receives a monthly token quota. You can monitor token usage for the current month or any of the previous six months in the API Consumption section.

Fig. A snapshot of the API Consumption report in LLM Usage Insights.

Only for AWS Bedrock (SearchUnify Partner Provisioned LLM, the usage bar is green when the token usage in a given month is significantly lower than the allocated quota. It turns blue as monthly usage approaches the limit, and it turns red once token consumption matches or exceeds the quota.

Fig. A snapshot of the API Consumption report in LLM Usage Insights.

You will receive email notifications when 50%, 90%, and 100% of the monthly token quota have been consumed. When token usage reaches 90% of the allocated quota, a CSM will reach out to you.

Fig. A snapshot of the email received when LLM token usage is hits 50%, 90%, and 100% of the monthly quota.