By ATS Staff - September 30th, 2026
FastAPI Infrastructure Latest Technologies PostgreSQL
Modern web applications need to handle increasing numbers of users, concurrent requests, large volumes of data, and resource-intensive operations without sacrificing performance or reliability. Applications built with FastAPI, Next.js, PostgreSQL, and Nginx provide a flexible foundation for developing scalable, high-performance web systems.
However, simply choosing these technologies does not guarantee that an application can handle heavy traffic. Performance depends on how each component is configured, how requests are processed, how database queries are optimized, and how server resources are managed.
An application running smoothly with 100 concurrent users may experience slow response times, database connection exhaustion, or even service interruptions when traffic increases to several thousand concurrent users.
This article explains how to design, configure, optimize, and maintain a FastAPI, Next.js, PostgreSQL, and Nginx application on an Ubuntu server to support heavy workloads. It covers server architecture, application servers, database optimization, caching, load balancing, security, monitoring, and practical deployment strategies.
A high-performance application should separate its responsibilities into distinct components. Each component should handle a specific part of the request-processing workflow.
Users and Clients
Web browsers, mobile applications, API clients
Nginx — Reverse Proxy
HTTPS, routing, static files, compression, rate limiting, load balancing
Next.js
Frontend, SSR, static rendering, UI
Node.js process
FastAPI
REST APIs, business logic, authentication
Uvicorn workers
PostgreSQL
Persistent storage, indexing, transactions, query processing
Optional Redis Cache and Background Workers
Caching, queues, asynchronous tasks, rate-limit counters
In a typical deployment:
For heavier workloads, these components can be distributed across multiple servers rather than running on a single Ubuntu machine.
The server's operating system and hardware form the foundation of application performance. Even well-optimized application code can struggle if the server has insufficient memory, slow storage, or inadequate CPU capacity.
The required resources depend on the application's workload, number of concurrent requests, database size, and processing requirements.
The following are illustrative starting configurations, not guaranteed capacity estimates.
| Workload | CPU | RAM | Storage |
|---|---|---|---|
| Development or testing | 2 vCPU | 4 GB | 40 GB SSD |
| Small production | 4 vCPU | 8 GB | 80 GB NVMe |
| Medium production | 8 vCPU | 16–32 GB | 160 GB NVMe |
| High-load production | 16+ vCPU | 32–64+ GB | High-performance NVMe |
For database-intensive applications, memory and storage performance are particularly important. For CPU-intensive workloads, such as data processing or complex calculations, CPU capacity can become the primary bottleneck.
It is also important to leave sufficient memory for Ubuntu, Nginx, monitoring, and other system services.
Start by updating Ubuntu and installing the basic tools required for server management.
sudo apt update
sudo apt upgrade -y
sudo apt install -y \
nginx \
postgresql \
postgresql-contrib \
python3 \
python3-venv \
python3-pip \
nodejs \
npm \
git \
curl \
ufw \
htop \
sysstat \
unzip
Install supported, compatible versions of Node.js and Python for the application's dependencies rather than assuming the Ubuntu repository version is appropriate for production.
Only expose the network ports required for public access and administration.
sudo ufw default deny incoming sudo ufw default allow outgoing sudo ufw allow OpenSSH sudo ufw allow 80/tcp sudo ufw allow 443/tcp sudo ufw enable sudo ufw status
PostgreSQL, FastAPI, and Next.js should normally listen on localhost or a private network rather than being directly accessible from the public internet.
Install and use monitoring utilities to understand server resource consumption.
htop
For memory:
free -h
For disk utilization:
df -h
For disk I/O:
iostat -xz 1
For system load:
uptime
Monitoring these metrics helps identify whether the application is CPU-bound, memory-constrained, or waiting on disk operations.
FastAPI is an asynchronous Python web framework designed for building APIs. It supports asynchronous request handling, dependency injection, data validation, and automatic API documentation.
However, FastAPI's asynchronous capabilities alone do not guarantee high throughput. The application must use appropriate concurrency, database connections, and resource management.
Uvicorn is an ASGI server that runs FastAPI applications. A single Uvicorn process can handle multiple concurrent asynchronous requests, but CPU-bound workloads and process-level isolation may benefit from multiple workers.
For production, run FastAPI using multiple worker processes, managed by a process manager such as systemd or another supported deployment mechanism.
Install the required packages:
pip install fastapi uvicorn[standard]
An example command using Uvicorn's worker support through Gunicorn is:
gunicorn main:app \
--worker-class uvicorn_worker.UvicornWorker \
--workers 4 \
--bind 127.0.0.1:8000 \
--timeout 60 \
--access-logfile -
This example assumes the uvicorn-worker package is installed:
pip install gunicorn uvicorn-worker
The number of workers should be based on CPU capacity, memory consumption, and load-testing results. Four workers may be suitable for a particular four-vCPU application, but it is not a universal optimal setting.
Each worker is a separate process and consumes its own memory and database connections.
FastAPI supports both synchronous and asynchronous endpoints.
An asynchronous endpoint is particularly useful when the request spends time waiting for external services, databases, or other I/O operations.
Example:
from fastapi import FastAPI
import httpx
app = FastAPI()
@app.get("/api/data")
async def get_data():
async with httpx.AsyncClient() as client:
response = await client.get(
"https://example.com/data",
timeout=10
)
return response.json()
For production applications, avoid creating a new HTTP client for every request. Instead, create and reuse a client through FastAPI's lifespan mechanism, allowing connection pooling and reducing repeated connection setup.
Use async def when the operations inside it are genuinely asynchronous. Calling blocking functions directly inside an asynchronous endpoint can block the event loop and reduce throughput.
For CPU-intensive operations, such as image processing, large calculations, or machine learning inference, consider a process pool or dedicated background workers.
Opening a new PostgreSQL connection for every API request is inefficient and can exhaust database resources.
Connection pooling allows multiple requests to reuse a limited number of database connections.
For example, with SQLAlchemy's asynchronous engine:
from sqlalchemy.ext.asyncio import (
create_async_engine,
async_sessionmaker
)
DATABASE_URL = (
"postgresql+asyncpg://appuser:password"
"@127.0.0.1:5432/appdb"
)
engine = create_async_engine(
DATABASE_URL,
pool_size=10,
max_overflow=5,
pool_timeout=30,
pool_recycle=1800,
pool_pre_ping=True
)
SessionLocal = async_sessionmaker(
engine,
expire_on_commit=False
)
The database session can then be managed through a FastAPI dependency:
from fastapi import Depends
from sqlalchemy.ext.asyncio import AsyncSession
async def get_db():
async with SessionLocal() as session:
yield session
And used in an endpoint:
@app.get("/api/customers")
async def get_customers(
db: AsyncSession = Depends(get_db)
):
result = await db.execute(
select(Customer).limit(100)
)
return result.scalars().all()
The example assumes that select, Customer, and SessionLocal are defined in the application.
The pool settings must be coordinated with the total number of application workers and the database's maximum connection capacity.
For example, four workers with a pool size of 10 and a maximum overflow of five could potentially request up to 60 connections in total. This needs to be considered alongside administrative connections and any other applications using PostgreSQL.
A common performance problem is executing synchronous, blocking operations inside asynchronous endpoints.
Examples include:
async def endpoints.Use asynchronous libraries for asynchronous workloads, or offload blocking operations to a thread pool or dedicated worker.
Returning thousands of database records in a single API response increases query time, memory usage, network traffic, and frontend rendering time.
Instead, implement pagination.
@app.get("/api/customers")
async def get_customers(
page: int = 1,
page_size: int = 25,
db: AsyncSession = Depends(get_db)
):
page = max(page, 1)
page_size = min(max(page_size, 1), 100)
offset = (page - 1) * page_size
result = await db.execute(
select(Customer)
.order_by(Customer.id)
.offset(offset)
.limit(page_size)
)
return result.scalars().all()
For very large tables, cursor-based pagination using indexed columns is often more efficient than increasingly large offsets.
Tasks such as generating PDF reports, sending bulk emails, processing large files, or running lengthy data calculations should not keep HTTP requests open unnecessarily.
A typical architecture uses:
The API can return a job identifier immediately, allowing the frontend to retrieve the status or results later.
FastAPI's built-in BackgroundTasks can be useful for lightweight tasks, but a durable external queue is generally more appropriate for long-running or critical work that must survive application restarts.
Next.js provides server-side rendering, static site generation, client-side navigation, and server components. How these features are used has a major impact on application performance.
Not every page needs to be rendered dynamically on every request.
| Rendering strategy | Suitable use cases | Performance considerations |
|---|---|---|
| Static generation | Marketing pages, documentation, articles | Low server processing for each visit |
| Incremental static regeneration | Frequently updated content | Regenerates pages when required |
| Server-side rendering | Personalized dashboards, dynamic content | Requires server processing |
| Client-side rendering | Highly interactive application views | Can reduce initial server rendering, but depends on client resources |
Static pages can be served directly by Nginx or a CDN, reducing the work performed by the Next.js server.
Dynamic rendering is useful for pages that depend on authentication, frequently changing information, or personalized content.
Choose the rendering approach based on actual page requirements rather than using server-side rendering for every page by default.
Next.js server components allow data to be fetched on the server without shipping the component's implementation to the browser.
For example:
export default async function CustomersPage() {
const response = await fetch(
"http://127.0.0.1:8000/api/customers",
{
cache: "no-store"
}
);
if (!response.ok) {
throw new Error("Unable to load customers");
}
const customers = await response.json();
return (
<main>
<h1>Customers</h1>
{customers.map((customer) => (
<div key={customer.id}>
{customer.name}
</div>
))}
</main>
);
}
This is an illustrative App Router example. The API URL and caching strategy should be adapted to the production architecture.
Avoid sequential requests when independent data can be retrieved concurrently:
const [customers, orders] = await Promise.all([
getCustomers(),
getOrders()
]);
This can reduce total page loading time when both requests are independent.
Large JavaScript bundles can increase initial page load time and browser CPU consumption.
To reduce unnecessary client-side processing:
For example, a complex charting component may be dynamically imported so it does not delay the initial rendering of a page where the chart is below the fold.
Large images can consume considerable bandwidth and slow down page rendering.
Next.js provides an image optimization component that supports responsive image sizing and modern image formats, depending on the configured loader and deployment environment.
import Image from "next/image";
export default function Banner() {
return (
<Image
src="/images/banner.webp"
alt="Application dashboard"
width={1200}
height={500}
priority
/>
);
}
Use priority only for important above-the-fold images, such as a primary hero image, rather than applying it to every image.
For applications with many images, a CDN or dedicated image service may further reduce bandwidth and processing demands.
Do not use the Next.js development server for production traffic.
Build the application:
npm run build
Run the production server:
npm run start
For a production deployment, manage the Next.js process through systemd or an appropriate process manager so that it restarts after unexpected failures and starts automatically after server reboots.
Where traffic is substantial, multiple Next.js instances can be run behind Nginx or a dedicated load balancer. Ensure that caching, session management, and any application state are compatible with multiple instances.
Nginx is a key component of the architecture. It handles incoming HTTP and HTTPS requests, routes traffic to the appropriate application, serves static content, and can distribute traffic across multiple application instances.
A typical application may run:
127.0.0.1:3000127.0.0.1:800080 and 443Nginx can route API requests to FastAPI and all other requests to Next.js.
Example Nginx configuration for a single application instance
upstream nextjs_backend {
server 127.0.0.1:3000;
keepalive 32;
}
upstream fastapi_backend {
server 127.0.0.1:8000;
keepalive 32;
}
server {
listen 80;
server_name example.com www.example.com;
return 301 https://example.com$request_uri;
}
server {
listen 443 ssl;
http2 on;
server_name example.com;
ssl_certificate
/etc/letsencrypt/live/example.com/fullchain.pem;
ssl_certificate_key
/etc/letsencrypt/live/example.com/privkey.pem;
client_max_body_size 20m;
keepalive_timeout 30;
sendfile on;
tcp_nopush on;
gzip on;
gzip_vary on;
gzip_min_length 1024;
gzip_types
application/json
application/javascript
application/xml
text/css
text/plain
image/svg+xml;
location /api/ {
proxy_pass http://fastapi_backend;
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For
$proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_connect_timeout 5s;
proxy_send_timeout 60s;
proxy_read_timeout 60s;
}
location / {
proxy_pass http://nextjs_backend;
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For
$proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_connect_timeout 5s;
proxy_send_timeout 60s;
proxy_read_timeout 60s;
}
}
This configuration assumes that the application routes API requests under /api/ and that FastAPI's routes match that prefix. The trailing slash in proxy_pass and the API route definitions should be adjusted together if URL prefixes need to be rewritten.
For applications that use WebSockets, add the appropriate Upgrade and Connection headers in a dedicated location block.
Nginx upstream keepalive connections reduce repeated TCP connection establishment between Nginx and the application servers.
The keepalive directive in the upstream configuration allows Nginx to retain idle upstream connections. The proxy configuration must also permit reuse, as shown above.
Connection reuse is especially useful when Nginx handles a large number of short API requests.
Nginx can handle many concurrent connections using its event-driven architecture.
A common baseline in /etc/nginx/nginx.conf is:
worker_processes auto;
events {
worker_connections 4096;
multi_accept on;
}
http {
sendfile on;
tcp_nopush on;
tcp_nodelay on;
keepalive_timeout 30;
include /etc/nginx/mime.types;
default_type application/octet-stream;
include /etc/nginx/conf.d/*.conf;
include /etc/nginx/sites-enabled/*;
}
worker_processes auto generally starts a worker for each available CPU core. worker_connections defines the maximum simultaneous connections per worker, subject to operating-system file descriptor limits and other constraints.
These settings do not represent the number of users the application can support. The actual capacity also depends on upstream connections, memory, application performance, and request duration.
Check the configuration before reloading:
sudo nginx -t sudo systemctl reload nginx
Gzip can reduce the size of text-based responses, including HTML, CSS, JavaScript, and JSON.
However, compressing already-compressed content, such as JPEG, WebP, or compressed video, generally provides little benefit and may waste CPU resources.
For very high traffic, evaluate whether compression should be handled by Nginx, a CDN, or both, and avoid redundant compression.
Rate limiting can protect APIs from excessive requests, automated abuse, and accidental traffic spikes.
For example, define a rate-limit zone in the http block:
limit_req_zone $binary_remote_addr
zone=api_limit:10m
rate=10r/s;
Apply it to a specific API location:
location /api/ {
limit_req zone=api_limit burst=20 nodelay;
proxy_pass http://fastapi_backend;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For
$proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
The example limits requests based on client IP, but real-world rate limits should reflect endpoint sensitivity, legitimate traffic patterns, and authenticated user identity where applicable.
If Nginx is behind a CDN or another proxy, configure trusted client IP handling correctly. Otherwise, multiple users may share a proxy IP, or clients may be able to spoof their addresses.
PostgreSQL is often one of the most important performance components of a data-intensive application. Poorly designed queries, missing indexes, excessive connections, and inefficient transactions can cause application-wide slowdowns.
PostgreSQL uses memory for caching, query operations, and maintaining database connections.
Important settings include:
| Setting | Purpose |
|---|---|
shared_buffers | Memory used for PostgreSQL's shared data cache |
effective_cache_size | Planner estimate of available cache |
work_mem | Memory available to individual query operations |
maintenance_work_mem | Memory used for maintenance tasks |
max_connections | Maximum permitted database connections |
An illustrative starting configuration for a dedicated PostgreSQL server with 16 GB RAM might look like:
shared_buffers = 4GB effective_cache_size = 12GB work_mem = 8MB maintenance_work_mem = 512MB max_connections = 100
These are not universal recommended values. For a shared application server, PostgreSQL may need substantially less memory to leave resources available for FastAPI, Next.js, and Nginx.
In particular, work_mem can be consumed by multiple query operations per connection, so setting it too high can result in excessive memory usage during concurrent workloads.
After modifying PostgreSQL settings, restart the service when required:
sudo systemctl restart postgresql
Indexes allow PostgreSQL to locate rows without scanning entire tables.
Consider a table containing millions of customer records:
CREATE TABLE customers (
id BIGSERIAL PRIMARY KEY,
first_name VARCHAR(100),
last_name VARCHAR(100),
email VARCHAR(255),
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);
If the application frequently searches customers by email, create an index:
CREATE UNIQUE INDEX idx_customers_email ON customers(email);
For queries that filter records by date:
CREATE INDEX idx_customers_created_at ON customers(created_at);
For queries that filter by status and sort by date, a composite index may be appropriate:
CREATE INDEX idx_orders_status_created ON orders(status, created_at DESC);
Indexes should reflect actual query patterns. Too many indexes increase storage usage and slow down inserts, updates, and deletes.
Use EXPLAIN ANALYZE to inspect query execution plans:
EXPLAIN ANALYZE SELECT * FROM customers WHERE email = 'customer@example.com';
Review execution time, rows scanned, index usage, and estimated versus actual row counts.
Selecting only the columns required by the application reduces data transfer, memory usage, and unnecessary processing.
Instead of:
SELECT * FROM customers;
Use:
SELECT id, first_name, last_name, email FROM customers LIMIT 100;
For large tables, combine selective queries with pagination and suitable indexes.
PostgreSQL is not designed to create unlimited simultaneous client connections.
An external connection pooler, such as PgBouncer, can help manage many application connections using a smaller number of database server connections.
Connection pooling architecture
FastAPI workers
Many concurrent API requests
PgBouncer
Queues and reuses database connections
PostgreSQL
Controlled number of database connections
PgBouncer supports different pooling modes, including session, transaction, and statement pooling. Transaction pooling can improve connection reuse, but some session-dependent features are incompatible with it. Check application requirements before selecting a mode.
Enable PostgreSQL's slow-query logging where appropriate:
log_min_duration_statement = 500
This example logs statements taking at least 500 milliseconds. The threshold should be chosen according to the application's performance targets and expected workload.
PostgreSQL's pg_stat_statements extension can help identify queries that consume the most total execution time or are executed frequently.
For example:
SELECT
query,
calls,
total_exec_time,
mean_exec_time
FROM pg_stat_statements
ORDER BY total_exec_time DESC
LIMIT 20;
This helps prioritize query optimization based on actual database workload rather than assumptions.
Transactions help maintain data consistency, but long-running transactions can hold locks, delay vacuum operations, and contribute to contention.
Keep transactions short, avoid waiting for external HTTP requests while holding database transactions open, and use suitable isolation levels for the application's requirements.
Regularly monitor deadlocks, lock waits, table bloat, and autovacuum activity as the database grows.
Redis can reduce repeated database queries and improve response times for frequently accessed data.
Without caching, every request for the same information may trigger another database query. With caching, the application can retrieve a previously computed result from memory.
FastAPI request
Redis cache
Is the requested data available?
Cache hit
Return cached data
Cache miss
Query PostgreSQL
PostgreSQL
Retrieve data and populate cache for future requests
On Ubuntu, Redis can be installed using:
sudo apt update sudo apt install redis-server
Check whether Redis is running:
sudo systemctl status redis-server
Test the connection:
redis-cli ping
A successful connection returns:
PONG
Redis should be configured to listen only on localhost or a secured private network, with authentication and other appropriate protections where needed.
Install the Redis Python client:
pip install redis
Example using the asynchronous Redis client:
import json
import redis.asyncio as redis
from fastapi import FastAPI
app = FastAPI()
cache = redis.Redis(
host="127.0.0.1",
port=6379,
decode_responses=True
)
@app.get("/api/dashboard")
async def get_dashboard():
cached = await cache.get("dashboard")
if cached:
return json.loads(cached)
data = {
"total_customers": 2500,
"total_orders": 8500
}
await cache.set(
"dashboard",
json.dumps(data),
ex=300
)
return data
This simplified example uses fixed data to demonstrate the caching mechanism. In a real application, the cache miss branch would retrieve the data from PostgreSQL.
The ex=300 parameter sets a five-minute time-to-live, after which the cached value expires.
Caching is particularly useful for:
Not all data should be cached. Financial transactions, sensitive user-specific data, and frequently changing records require careful cache design.
Cache invalidation is also important. When data changes, the application should either invalidate the relevant cache entry or update it. Otherwise, users may receive outdated information.
For personalized responses, include the appropriate user or tenant identifier in the cache key to prevent data from being shared across users.
Some workloads do not need to be completed while the user is waiting for an HTTP response.
Examples include:
Instead of processing these tasks directly inside an API request, place them in a background queue.
Next.js frontend
FastAPI
Accept request and enqueue job
Redis / Message Queue
Store pending jobs
Worker 1
Worker 2
Workers process jobs independently and store their results or status.
Tools such as Celery and RQ can manage background tasks using Redis or compatible brokers.
For reliable processing, design jobs to be retryable and idempotent where possible. Track job states, failures, retry counts, and completion times. This prevents temporary service interruptions from resulting in lost or duplicated business operations.
When one server can no longer meet the application's performance targets, adding more resources to the same server may help, but eventually horizontal scaling becomes necessary.
Horizontal scaling means running multiple instances of an application and distributing incoming requests among them.
Internet traffic
Ubuntu Server
Nginx
Next.js
FastAPI
PostgreSQL
Redis
A single-server architecture is relatively straightforward to deploy and maintain. It can be sufficient for many small and medium applications, depending on actual workload.
However, it introduces a single point of failure. If the server goes offline, the entire application may become unavailable.
Internet traffic
Load Balancer / Nginx
TLS termination, routing, health checks
Application Server 1
Next.js
FastAPI
Application Server 2
Next.js
FastAPI
PostgreSQL Server
Redis Server
A multi-server deployment separates application processing from database storage and other shared services.
This allows application instances to scale independently, while centralized data services maintain shared application state.
Important considerations include:
Nginx can distribute requests among multiple upstream servers using round-robin, least-connections, or other supported methods. For larger deployments, a managed load balancer or dedicated load-balancing tier may be appropriate.
Database scaling requires particular attention because it contains shared application state.
Potential approaches include:
Read replicas can help distribute read queries, but replication lag means they may not immediately reflect recent writes. Transactions and read-after-write operations may need to be routed to the primary database.
Security is essential for a high-load production application. A compromised server or resource-exhaustion attack can cause outages regardless of the application's performance optimizations.
FastAPI applications should implement appropriate authentication and authorization, input validation, request-size limits, and protection against abusive traffic.
Consider:
Do not rely exclusively on frontend checks to enforce access control. All sensitive permissions must be validated by the backend.
Create a dedicated database user for the application instead of using the PostgreSQL superuser.
For example:
CREATE USER appuser WITH PASSWORD 'replace-with-a-strong-secret'; GRANT CONNECT ON DATABASE appdb TO appuser;
Connect to appdb as a database administrator, then grant only the necessary schema and table privileges:
GRANT USAGE ON SCHEMA public TO appuser; GRANT SELECT, INSERT, UPDATE, DELETE ON ALL TABLES IN SCHEMA public TO appuser; GRANT USAGE, SELECT ON ALL SEQUENCES IN SCHEMA public TO appuser;
For future tables and sequences, configure default privileges for the role that creates them:
ALTER DEFAULT PRIVILEGES FOR ROLE dbowner IN SCHEMA public GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO appuser; ALTER DEFAULT PRIVILEGES FOR ROLE dbowner IN SCHEMA public GRANT USAGE, SELECT ON SEQUENCES TO appuser;
Replace dbowner with the actual role used to create application objects. Restrict access further for applications that only need read or limited write permissions.
Rate limiting at Nginx, application-level quotas, and suitable CDN or web application firewall protections can help control abusive requests.
For large public applications, a CDN may also provide distributed traffic absorption, caching, and additional security features.
Rate limiting should not replace proper capacity planning or infrastructure monitoring.
A production application needs continuous monitoring to identify performance bottlenecks before they cause significant service interruptions.
Monitoring should cover the operating system, Nginx, FastAPI, Next.js, PostgreSQL, and Redis.
| Tool | Purpose |
|---|---|
| Prometheus | Collects time-series metrics |
| Grafana | Visualizes metrics and dashboards |
| Netdata | Real-time server monitoring |
| Loki | Centralized log collection |
| Sentry | Tracks application errors |
| pg_stat_statements | Identifies expensive database queries |
Recommended metrics
| Metric | What it indicates |
| CPU utilization | Application processing pressure and CPU saturation |
| Memory utilization | Memory pressure, cache usage, and potential out-of-memory conditions |
| Disk I/O latency | Storage bottlenecks and database I/O performance |
| API response time | How quickly API requests are processed |
| Requests per second | Application throughput |
| HTTP error rate | Application failures and upstream issues |
| Database connections | Connection usage and potential exhaustion |
| Cache hit ratio | How effectively cached data is being reused |
| Queue depth | Whether background workers are keeping up with incoming jobs |
Monitor latency percentiles, particularly p95 and p99, rather than relying only on average response time. A low average can conceal slow responses affecting a significant portion of users.
Define service-level objectives (SLOs) for response times, availability, and error rates based on the application's actual requirements.
Load testing helps establish how much traffic the application can handle and where performance begins to degrade.
Testing should be performed in an isolated staging environment or with controlled production testing that will not disrupt real users.
Popular tools include:
A simple load-testing script can simulate concurrent requests to a FastAPI endpoint.
import http from "k6/http";
import { check, sleep } from "k6";
export const options = {
stages: [
{ duration: "1m", target: 50 },
{ duration: "3m", target: 100 },
{ duration: "3m", target: 250 },
{ duration: "1m", target: 0 }
],
thresholds: {
http_req_failed: ["rate<0.01"],
http_req_duration: ["p(95)<500"]
}
};
export default function () {
const response = http.get(
"https://example.com/api/customers"
);
check(response, {
"status is 200": (r) => r.status === 200
});
sleep(1);
}
Run the test with:
k6 run load-test.js
This example ramps up to 250 virtual users. Virtual users are not the same as requests per second or actual simultaneous server-side requests, so interpret the results in the context of the script's request rate and response times.
The sample thresholds specify a target of less than 1% failed requests and a p95 response time below 500 ms. They are illustrative targets, not requirements for every application.
Suggested load-testing stages
During every test, monitor application response times, CPU utilization, memory, PostgreSQL query latency, database connections, and Nginx error logs.
The purpose is not simply to reach a large number of users, but to understand the conditions under which the application continues to meet its performance and reliability requirements.
On an Ubuntu server, systemd can manage Next.js and FastAPI as persistent services. This ensures they start on reboot and can be restarted when they fail.
Create a systemd unit:
sudo nano /etc/systemd/system/fastapi.service
Example:
[Unit]
Description=FastAPI Application
After=network.target postgresql.service
[Service]
User=www-data
Group=www-data
WorkingDirectory=/var/www/myapp/backend
EnvironmentFile=/etc/myapp/backend.env
ExecStart=/var/www/myapp/backend/venv/bin/gunicorn \
main:app \
--worker-class uvicorn_worker.UvicornWorker \
--workers 4 \
--bind 127.0.0.1:8000 \
--timeout 60
Restart=always
RestartSec=5
NoNewPrivileges=true
PrivateTmp=true
[Install]
WantedBy=multi-user.target
Ensure that www-data has the necessary read and execute permissions for the application directory and appropriate access to its environment file. Adjust the worker count based on testing.
Activate the service:
sudo systemctl daemon-reload sudo systemctl enable fastapi sudo systemctl start fastapi sudo systemctl status fastapi
Create another service:
sudo nano /etc/systemd/system/nextjs.service
Example:
[Unit] Description=Next.js Application After=network.target [Service] Type=simple User=www-data Group=www-data WorkingDirectory=/var/www/myapp/frontend Environment=NODE_ENV=production Environment=PORT=3000 EnvironmentFile=/etc/myapp/frontend.env ExecStart=/usr/bin/npm run start Restart=always RestartSec=5 NoNewPrivileges=true PrivateTmp=true [Install] WantedBy=multi-user.target
Confirm the correct path to the Node.js and npm executables using:
which node which npm
Then enable and start the service:
sudo systemctl daemon-reload sudo systemctl enable nextjs sudo systemctl start nextjs sudo systemctl status nextjs
For production, use a consistent Node.js installation path and ensure the application has been built before starting the service.
For FastAPI:
sudo journalctl -u fastapi -f
For Next.js:
sudo journalctl -u nextjs -f
For Nginx:
sudo tail -f /var/log/nginx/access.log sudo tail -f /var/log/nginx/error.log
For PostgreSQL:
sudo journalctl -u postgresql -f
Centralized log collection becomes increasingly useful when the application runs across multiple servers.
Performance and scalability are only part of a reliable production deployment. Data protection and recovery planning are equally important.
A basic PostgreSQL logical backup can be created using:
pg_dump -U appuser -d appdb \
-F c \
-f /backup/appdb.dump
Restore it using:
pg_restore \
-U appuser \
-d appdb \
/backup/appdb.dump
For larger databases or more stringent recovery requirements, consider physical backups and continuous archiving with PostgreSQL's write-ahead logs (WAL).
Maintain:
Backups should be encrypted where appropriate, access-controlled, and protected from accidental deletion or ransomware.
A backup is useful only if it can be restored successfully.
Periodically restore backups to an isolated environment and verify database consistency, application functionality, and recovery time.
Define:
These requirements should influence the backup frequency, replication strategy, and disaster recovery architecture.
For an application that is expected to experience substantial traffic, the following is an example of a distributed deployment.
Users and Internet Traffic
Web browsers, mobile apps, external API clients
CDN / Load Balancer
TLS, caching, traffic distribution, security controls
App Server 1
Nginx
Next.js
FastAPI
App Server 2
Nginx
Next.js
FastAPI
PostgreSQL
Primary database and optional read replicas
Redis
Caching and queue infrastructure
Background Workers
Independent processing for long-running jobs
Centralized Monitoring and Logging
Metrics, alerts, tracing, logs, and operational dashboards
This architecture allows different parts of the application to scale independently. For example, more application servers can be added when API traffic increases, while background worker capacity can be increased separately when processing queues grow.
The database should be scaled based on query workload and data access patterns rather than simply matching the number of application servers.
Production readiness
0 of 14 completeUse production builds for Next.js and a production ASGI server for FastAPIConfigure Nginx as a reverse proxy with HTTPS and connection reuseUse asynchronous I/O for suitable FastAPI endpointsImplement database connection pooling and sensible connection limitsAdd indexes based on real query patterns and inspect slow queriesUse pagination and limit unnecessary data retrievalApply suitable Next.js rendering strategies and reduce client-side JavaScriptIntroduce Redis caching for frequently accessed data where appropriateMove lengthy tasks to background workersConfigure rate limiting and secure all public-facing endpointsMonitor CPU, memory, disk, database, and application metricsRun load, stress, spike, and soak tests in a controlled environmentSet up automated backups and regularly test restorationEstablish performance objectives, alerting, and incident response proceduresReset checklist
Building a high-performance FastAPI, Next.js, PostgreSQL, and Nginx application on Ubuntu requires a combination of efficient application design, optimized database operations, reliable infrastructure, and continuous monitoring.
FastAPI provides asynchronous API processing, Next.js supports multiple rendering strategies for fast web experiences, PostgreSQL offers robust relational data management, and Nginx provides efficient request routing and traffic handling.
Redis caching, database connection pooling, background processing, and horizontal scaling can further improve performance when they address measured bottlenecks.
The most important principle is to optimize based on actual workload measurements rather than assumptions. Begin with a well-configured single-server deployment when appropriate, establish performance benchmarks, identify bottlenecks, and scale individual components as the application's traffic and operational requirements grow.
With suitable architecture, testing, monitoring, and recovery procedures, this technology stack can support demanding production workloads while remaining maintainable and adaptable as the application expands.