Quickstart¶
The repository includes runnable example assets in examples/. Generated run artifacts are written under examples/runs/ and are intentionally ignored by Git.
Single-item example¶
The bundled single-item example uses query_id as the canonical row id, returns a bare JSON object, and lets the parser infer that id from the request context.
Submit the example run and poll once:
export OPEN_AI_KEY="sk-..."
PYTHONPATH=src .venv/bin/python -m llm_batch_annotate.cli run examples/config/run_config.json --run-id example-single --no-poll-until-terminal
Resume the same run later and wait up to 10 polls, every 2 minutes:
PYTHONPATH=src .venv/bin/python -m llm_batch_annotate.cli resume examples/config/run_config.json example-single --poll-interval 2m --max-polls 10
Grouped example¶
The grouped example sends up to 3 units per request, so the bundled 10-row CSV becomes 4 requests with sizes 3, 3, 3, 1.
export OPEN_AI_KEY="sk-..."
PYTHONPATH=src .venv/bin/python -m llm_batch_annotate.cli run examples/config/run_config_2.json --run-id example-grouped --no-poll-until-terminal
Resume and continue polling:
PYTHONPATH=src .venv/bin/python -m llm_batch_annotate.cli resume examples/config/run_config_2.json example-grouped --poll-interval 30s --max-polls 20
Artifacts to inspect¶
After a run finishes, look at:
metadata/manifest.jsonfor the lifecycle state and handle metadata.metadata/summary.jsonfor the CLI-friendly summary snapshot.raw/raw_outputs.jsonlandraw/raw_errors.jsonlfor provider results.parsed/parsed_requests.jsonlfor request-level parsed payloads.parsed/responses.jsonlfor per-row outputs, withquery_idat top level for direct merge-back to the source table.parsed/failures.jsonlfor parser, provider, flattening, or validation failures.