Skip to content

Python / reports

Report PDFs in Python

Post your report HTML to a render API from an async endpoint that returns a streaming response, then stream the response with httpx.stream and yield chunks. No browser binary in your Python deployment.

Why reports become PDFs

A report is a snapshot of numbers at a moment in time. Its whole value is that it does not change when the underlying dashboard does. That is the difference between a document and a view of a database, and it is why analytics, operations and account teams keep asking for this.

What a report has to carry

Before any of the code below matters, the template has to produce a document that is actually a report. These are the fields that make it one, and the ones a reader or an auditor will look for first.

  • the reporting period as an explicit start and end date
  • the generation timestamp, distinct from the period
  • the data source or filter set, so a reader can tell what was excluded
  • each metric with its unit, since a bare number is not a figure
  • comparison against the prior period where the report claims a trend

The Python implementation

httpx is the client most projects here already have. The render happens in an async endpoint that returns a streaming response, and you stream the response with httpx.stream and yield chunks.

python

import os, httpx
from fastapi import FastAPI
from fastapi.responses import StreamingResponse

app = FastAPI()

@app.get("/report/{report_id}.pdf")
async def report_pdf(report_id: str):
    html = render_report(report_id)

    client = httpx.AsyncClient(timeout=60.0)
    request = client.build_request(
        "POST",
        "https://api.pdfpipe.xyz/v1/pdf",
        headers={"Authorization": f"Bearer {os.environ['PDFPIPE_KEY']}"},
        json={"html": html, "options": {"format": "A4", "printBackground": True}},
    )
    response = await client.send(request, stream=True)

    # stream=True keeps a long report off the worker's heap.
    return StreamingResponse(
        response.aiter_bytes(),
        media_type="application/pdf",
        background=response.aclose,
    )

The CSS that makes a report page correctly

The layout problem specific to this document is that charts and tables must not split mid-element, and a report long enough to matter will page-break somewhere awkward unless break rules are set. These rules handle it.

css

/* Charts and tables must not split. Sections start on a fresh
   page so a heading never ends up alone at the foot of one. */
.chart, figure, table { break-inside: avoid; }
h2 { break-after: avoid; }
section { break-before: page; }
section:first-of-type { break-before: auto; }

@page { size: A4; margin: 20mm 18mm; }

What goes wrong in Python

requests holds the entire body in memory before you touch it. httpx.stream does not, and on documents past a few megabytes that difference decides whether your worker survives a burst.

What people try first

Most Python projects reach for WeasyPrint, which is excellent at CSS Paged Media but needs Cairo, Pango and GDK-PixBuf present in the image, and diverges from a browser on modern layout. That works until it is running on more than one machine, at which point the browser becomes the thing you operate rather than the thing you use.

Where the report lives afterwards

Rendering is the short part. Reports pile up faster than any other document here because they are produced on a schedule whether or not anyone reads them. Most teams keep twelve months and expire the rest, which makes storage lifecycle a design decision rather than an afterthought.

Getting the document right

  • Check the period the figures cover, because a report without an unambiguous date range is worse than no report.
  • Generation is triggered by a schedule, usually month end or week end, so size the timeout for that path rather than for a health check.
  • This is a batch shape: render the run through the batch endpoint rather than firing thousands of individual requests.
  • Backgrounds are painted by default here, so a design that uses colour needs nothing set. Only an explicit print_background of false turns them off.

Frequently asked

Do I need Chromium installed to generate reports from Python?

No. The render happens over HTTP, so your Python deployment stays the size it is now. That is the main reason to use an API rather than WeasyPrint, which is excellent at CSS Paged Media but needs Cairo, Pango and GDK-PixBuf present in the image, and diverges from a browser on modern layout.

How do I stop a long report using all the memory?

Stream the response with httpx.stream and yield chunks. requests holds the entire body in memory before you touch it. httpx.stream does not, and on documents past a few megabytes that difference decides whether your worker survives a burst.

What has to be on a report?

At minimum: the reporting period as an explicit start and end date; the generation timestamp, distinct from the period; the data source or filter set, so a reader can tell what was excluded. The one to get right before anything else is the period the figures cover, because a report without an unambiguous date range is worse than no report.

Do I need to store the generated reports?

Depends on the document, and this one has a clear answer: reports pile up faster than any other document here because they are produced on a schedule whether or not anyone reads them. Most teams keep twelve months and expire the rest, which makes storage lifecycle a design decision rather than an afterthought.

Can I keep my existing report template?

Yes, if it produces HTML. Whatever renders your report view today can render the same markup for the PDF, which is why the CSS above is the only new thing you write.

Related

Other Python documents, and the same report in other stacks.

Certificate PDFs in Pythona certificate is meant to be shown to a third party who has no relationship with the issuing systemStatement PDFs in Pythona statement summarises a period that is now closed. Regenerating it later from live data would produce a different documentQuote PDFs in Pythona quote is an offer with an expiryPurchase order PDFs in Pythona purchase order is the document a supplier's accounts team matches an invoice againstPayslip PDFs in Pythona payslip is a personal financial record an employee keepsReport PDFs in PHPUsing cURL, in a controller action that echoes the body with a PDF Content-Type.Report PDFs in RubyUsing Net::HTTP, in a controller action using send_data.Report PDFs in GoUsing net/http, in an http.HandlerFunc that copies the body through.Report PDFs in JavaUsing java.net.http.HttpClient, in a controller returning StreamingResponseBody.Report PDFs in RustUsing reqwest, in an axum handler returning a streaming Body.When the output is wrongBlank pages, missing backgrounds, breaks in the wrong place, by symptom.All stacks and documentsThe full grid of what this covers.the full Python guideSetup, error handling and the deployment shape, not just the request.report generation in productionVolume, storage and the delivery step, for teams already shipping these.how this compares to weasyprintThe tradeoff written out, including where the incumbent is the better call.numbers, dates and currencyFormatting that has to be right on a document somebody keeps.

100 free documents a month, no card.