server side tools, rate limiting, connections, logprobs, token usage, configurable models knowt ·

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/32

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 11:02 PM on 8/3/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

33 Terms

1
New cards
2
New cards
3
New cards
4
New cards
5
New cards
What is server-side tool use in the context of chat models?
Some providers support server-side tool-calling loops where models can interact with web search; code interpreters; and other tools and analyze results within a single conversational turn.
6
New cards
What happens to the response message content when a model calls a tool on the server side?
It will include content representing the invocation and result of the tool; returned via content blocks in a provider-agnostic format.
7
New cards
What parameter do chat model integrations accept to help manage rate limits?
The `rate_limiter` parameter; which can be provided during initialization to control the rate at which requests are made.
8
New cards
What built-in rate limiter does LangChain provide; and what is a key property of it?
The `InMemoryRateLimiter`; which is thread safe and can be shared by multiple threads in the same process.
9
New cards
What is the `base_url` parameter used for?
It specifies the server address LangChain talks to instead of the default (e.g. api.openai.com); used for your own server such as vLLM or Together AI.
10
New cards
Why can't you just set `base_url` to connect to OpenRouter or LiteLLM?
Because they are not just custom URLs but dedicated services with their own features (routing; fallback models; caching); using plain `base_url` loses the provider's features; so you need dedicated classes like `ChatOpenRouter` or `ChatLiteLLM`.
11
New cards
What is an HTTP Proxy used for in this context?
It's an intermediary between your code and the internet; in corporate networks; you can't reach an external API without one.
12
New cards
What parameter is used to set a corporate HTTP proxy on `ChatOpenAI`?
`openai_proxy`.
13
New cards
How do you enable log probabilities on a model?
By setting the `logprobs` parameter when initializing the model.
14
New cards
Where is token usage information included when available from a provider?
On the `AIMessage` objects produced by the corresponding model.
15
New cards
What special requirement do OpenAI and Azure OpenAI chat completions have regarding token usage data in streaming?
They require users to explicitly opt-in to receiving token usage data in streaming contexts.
16
New cards
How can you track aggregate token counts across multiple models in an application?
Using either a callback or a context manager.
17
New cards
What does the `config` parameter passed to `invoke` use to provide run-time control?
A `RunnableConfig` dictionary; which provides run-time control over execution behavior; callbacks; and metadata tracking.
18
New cards
What are common configuration options passed via `config`?
`run_name`; `tags`; `metadata`; `callbacks`.
19
New cards
What is a callback; as described in the notes?
'When X happens — do Y.'
20
New cards
What is the limitation of a model configured without configurable fields; as described in the notes?
The model is fixed forever — e.g. `model = ChatOpenAI(model=\gpt-4\')` — and switching to a different model requires rewriting the code.'
21
New cards
How do you make a model's `model` field configurable at runtime instead of fixed?
Don't specify `model` at initialization — it becomes configurable; and the same object can then serve different models depending on `config={\configurable\': {\'model\': \'...\'}}`.'
22
New cards
Why is `config_prefix` needed when a chain has multiple configurable models?
To disambiguate which configuration setting applies to which model.
23
New cards
How does the 'Declarative Operations' approach to configurable models work?
Tools are bound once via `bind_tools`; while the underlying model changes via `config` on each call.
24
New cards
How does 'Dynamic Model Selection' via middleware work?
The model is selected automatically by logic; e.g. based on conversation length — a simple model for short dialogues; an advanced one for long ones.
25
New cards
According to the comparison table; which configurable-model approach has the highest flexibility rating?
Dynamic middleware (5 stars); compared to Configurable fields and Declarative + config (3 stars each).
26
New cards
What use case is suggested for configurable models when comparing GPT and Claude in production without code changes?
A/B testing — changing the model via config without touching the code.
27
New cards
What use case is suggested for configurable models to reduce costs?
Cost optimization — routing simple tasks to a cheap model and complex tasks to an expensive model.
28
New cards
What use case is suggested for configurable models if a primary provider goes down?
Fallback — e.g. if GPT goes down; switch to Claude.
29
New cards
python from langchain.rate_limiters import InMemoryRateLimiter rate_limiter = InMemoryRateLimiter( ________=0.1; # 1 request every 10s check_every_n_seconds=0.1; # Check every 100ms whether allowed to make a request ________=10; # Controls the maximum burst size. ) model = init_chat_model( model='gpt-5.5'; model_provider='openai'; rate_limiter=rate_limiter )
```python from langchain.rate_limiters import InMemoryRateLimiter rate_limiter = InMemoryRateLimiter( requests_per_second=0.1; # 1 request every 10s check_every_n_seconds=0.1; # Check every 100ms whether allowed to make a request max_bucket_size=10; # Controls the maximum burst size. ) model = init_chat_model( model='gpt-5.5'; model_provider='openai'; rate_limiter=rate_limiter ) ```
30
New cards
python from langchain_openai import ChatOpenAI model = ________( model='llama-3-8b'; ________='https://api.together.xyz/v1'; # ← custom server api_key='YOUR_KEY' )
```python from langchain_openai import ChatOpenAI model = ChatOpenAI( model='llama-3-8b'; base_url='https://api.together.xyz/v1'; # ← custom server api_key='YOUR_KEY' ) ```
31
New cards
python # ✅ Correct — use dedicated integration classes from langchain_openrouter import ________ model = ChatOpenRouter(model='gpt-4') from langchain_litellm import ________ model = ChatLiteLLM(model='gpt-4')
```python # ✅ Correct — use dedicated integration classes from langchain_openrouter import ChatOpenRouter model = ChatOpenRouter(model='gpt-4') from langchain_litellm import ChatLiteLLM model = ChatLiteLLM(model='gpt-4') ```
32
New cards
python from langchain.chat_models import init_chat_model from langchain_core.callbacks import ________ model_1 = init_chat_model(model='gpt-5.4-mini') model_2 = init_chat_model(model='claude-haiku-4-5-20251001') callback = UsageMetadataCallbackHandler() result_1 = model_1.invoke('Hello'; config={'callbacks': [callback]}) result_2 = model_2.invoke('Hello'; config={'callbacks': [callback]}) print(callback.________)
```python from langchain.chat_models import init_chat_model from langchain_core.callbacks import UsageMetadataCallbackHandler model_1 = init_chat_model(model='gpt-5.4-mini') model_2 = init_chat_model(model='claude-haiku-4-5-20251001') callback = UsageMetadataCallbackHandler() result_1 = model_1.invoke('Hello'; config={'callbacks': [callback]}) result_2 = model_2.invoke('Hello'; config={'callbacks': [callback]}) print(callback.usage_metadata) ```
33
New cards
python from langchain.agents.middleware import ________; ModelRequest; ModelResponse basic_model = ChatOpenAI(model='gpt-5.4-mini') # cheap advanced_model = ChatOpenAI(model='gpt-5.5') # expensive but smarter @wrap_model_call def dynamic_model_selection(request: ModelRequest; handler) -> ModelResponse: message_count = len(request.state['messages']) if message_count > 10: model = advanced_model else: model = basic_model return handler(request.________(model=model)) agent = create_agent( model=basic_model; # default tools=tools; middleware=[dynamic_model_selection]; # ← our middleware )
```python from langchain.agents.middleware import wrap_model_call; ModelRequest; ModelResponse basic_model = ChatOpenAI(model='gpt-5.4-mini') # cheap advanced_model = ChatOpenAI(model='gpt-5.5') # expensive but smarter @wrap_model_call def dynamic_model_selection(request: ModelRequest; handler) -> ModelResponse: message_count = len(request.state['messages']) if message_count > 10: model = advanced_model else: model = basic_model return handler(request.override(model=model)) agent = create_agent( model=basic_model; # default tools=tools; middleware=[dynamic_model_selection]; # ← our middleware ) ```