Files on object storage: ObjectStorageToolset¶
Note
Experimental: this can change or be removed in a minor release of this provider. See Stable and experimental features.
Give an agent the files under one location in S3, GCS, Azure Blob Storage or any other
store Airflow’s ObjectStoragePath can open, and let it find and
read what it needs: a month’s reports, the config files of a failing job, the first
rows of a Parquet extract. The agent can only read, and only under the path you give
it.
@dag(tags=["example"])
def example_object_storage_toolset():
@task.agent(
llm_conn_id="pydanticai_default",
system_prompt=(
"You answer questions about the files in a reports bucket. List directories "
"to find what exists, then read the files you need. Cite the files you used."
),
toolsets=[
# Read-only, and every path the model names is resolved under this root.
ObjectStorageToolset("s3://acme-reports/finance/", conn_id="aws_default"),
],
)
def summarize_month(month: str) -> str:
return f"Summarize revenue and the biggest changes for {month}."
summarize_month("2026-09")
The credentials come from conn_id, the same connection your other tasks use for
that store. With conn_id=None the store’s default credentials apply, such as the
worker’s AWS role.
The three tools¶
list_filesLists one directory, sorted by name: files with their size in bytes, and subdirectories with a trailing slash. The model goes deeper by listing a subdirectory. A directory with more than
max_filesentries is listed a page at a time, with a note giving theoffsetto continue from. Every call lists the whole directory from storage, so a very large one is slow to page through.get_file_infoThe size and last-modified time of one file, so the model can decide whether it is worth reading.
read_fileA text file a window of lines at a time. The model passes
offsetandlimitas line numbers, and a result that stops early says whichoffsetto continue from, the same shape as the sandbox’sread_file. A Parquet or Avro file comes back as its row count, its schema and its first 20 rows;offsetandlimitdo not apply to it. To query one, such as summing a column, use Files with DataFusion: DataFusionToolset. Text compressed as.gz,.bz2or.xzis decompressed first.
Every path the model supplies is relative to the root. An absolute path, a path with a
scheme such as s3://, and a path that climbs out with .. are refused, and on a
local root a symbolic link that points outside the root is refused too, and left out of
listings.
What the model cannot read is refused with a message it can act on, rather than failing
the task: an image or PDF, a binary file, a file larger than max_read_bytes (10 MiB
by default), a corrupt file, a path that does not exist, or one the connection may not
read.
Restricting the agent¶
This agent can read only under s3://acme-reports/finance/, and every list and read
is bounded:
@dag(tags=["example"])
def example_object_storage_toolset_restricted():
AgentOperator(
task_id="read_finance_reports",
prompt="Summarize the September finance report.",
llm_conn_id="pydanticai_default",
toolsets=[
ObjectStorageToolset(
# Resolve every path the model names under this root. The toolset never writes.
"s3://acme-reports/finance/",
# Credentials scoped to that prefix keep the limit if the path check has a gap.
conn_id="aws_reports_reader",
# List at most 50 entries per call; the model pages through the rest.
max_files=50,
# Refuse a file larger than 1 MiB, measured after decompression.
max_read_bytes=1024 * 1024,
# Return at most 16 KiB from one read; the model pages through a longer text file.
max_output_bytes=16 * 1024,
# Name the tools reports_list_files, reports_get_file_info and reports_read_file.
tool_prefix="reports",
# Let the model correct an invalid call up to 2 times.
max_retries=2,
)
],
)
Run against an S3 endpoint where an acme-payroll bucket sits beside the reports,
these reads were refused. Each refusal comes back to the model as the tool’s result,
so it does not use max_retries, and the run carries on:
../../acme-payroll/salaries.csv'../../acme-payroll/salaries.csv' leaves the storage root; '..' is not allowed.s3://acme-payroll/salaries.csv's3://acme-payroll/salaries.csv' is not a relative path. Name files relative to the storage root./etc/passwd'/etc/passwd' is not a relative path. Name files relative to the storage root.2026-09/chart.png'2026-09/chart.png' is a png file, which this tool cannot read as text.2026-09/ledger.csv, a 1.5 MB file'2026-09/ledger.csv' is larger than the 1.0MB this tool reads.
Because refused reads cost nothing from max_retries, bound a run that keeps
asking with the operator’s usage_limits. The path check is the toolset’s only
boundary on location; credentials that can read only s3://acme-reports/finance/ keep that limit
if the check has a gap.
Parameters¶
pathThe root, for example
"s3://acme-reports/finance/". Templated when the toolset is passed toAgentOperatoror@task.agent.conn_idThe connection for the store. Templated like
path.max_filesThe most entries one
list_filesresult holds. Default200.max_read_bytesThe largest file
read_fileopens, after decompression. Default 10 MiB.max_output_bytesThe most bytes one
read_fileresult holds. Default 50 KiB.tool_prefixA prefix for the tool names, needed when the agent has another toolset with the same tool names: a second
ObjectStorageToolset, or aSandboxToolset, which has aread_fileof its own.max_retriesHow many times the model may correct a call with invalid arguments. A failed read does not count. Default
None, the agent’sretries. See How often the model may correct a failed call.
When to choose it¶
Choose it when the agent should read files the way a person would: browse a
directory, open the report that matters, look at a file’s first lines or rows. Text
files need nothing beyond the filesystem package for your store, which Airflow’s object
storage already uses; Parquet and Avro need this provider’s parquet or avro
extra. To hand the model a known file rather than let it find one, @task.llm_file_analysis
reads the file for it. For read-only access through a hook’s own methods,
HookToolset(S3Hook(), allowed_methods=["list_keys", "read_key"]) works too, but the
model then chooses the bucket and key. Pinning bucket_name with pinned_arguments
fixes the bucket and still leaves the model any key in it, where this toolset keeps the
model under one root.
What it cannot do
It cannot write, delete, move or copy anything.
It does not search inside files. The model finds a file by listing directories, so a store with thousands of files per directory is slow to explore. Give it a narrower root.
It reads one file at a time. To ask a question across many files, such as a sum over a month of Parquet files, use Files with DataFusion: DataFusionToolset, which runs SQL over them.
The root is the boundary only as far as the path check goes. The connection’s own permissions are the real limit on what can be read, so scope its role or key to the prefix you pass as
path.
Using it with other agent frameworks¶
ObjectStorageToolset implements
ToolProvider, so the same three tools work in
a Strands or Google ADK agent through AirflowTools; see Agent frameworks.
Outside AgentOperator, path and conn_id are used as given: they are not
rendered as templates.