Build a High-Performance FastAPI, Next.js, PostgreSQL, and Nginx Application for Heavy Loads on Ubuntu Server



By ATS Staff - September 30th, 2026

FastAPI   Infrastructure  Latest Technologies  PostgreSQL  

Introduction

Modern web applications need to handle increasing numbers of users, concurrent requests, large volumes of data, and resource-intensive operations without sacrificing performance or reliability. Applications built with FastAPI, Next.js, PostgreSQL, and Nginx provide a flexible foundation for developing scalable, high-performance web systems.

However, simply choosing these technologies does not guarantee that an application can handle heavy traffic. Performance depends on how each component is configured, how requests are processed, how database queries are optimized, and how server resources are managed.

An application running smoothly with 100 concurrent users may experience slow response times, database connection exhaustion, or even service interruptions when traffic increases to several thousand concurrent users.

This article explains how to design, configure, optimize, and maintain a FastAPI, Next.js, PostgreSQL, and Nginx application on an Ubuntu server to support heavy workloads. It covers server architecture, application servers, database optimization, caching, load balancing, security, monitoring, and practical deployment strategies.

1. Understanding the Application Architecture

A high-performance application should separate its responsibilities into distinct components. Each component should handle a specific part of the request-processing workflow.

Users and Clients

Web browsers, mobile applications, API clients

Nginx — Reverse Proxy

HTTPS, routing, static files, compression, rate limiting, load balancing

Next.js

Frontend, SSR, static rendering, UI

Node.js process

FastAPI

REST APIs, business logic, authentication

Uvicorn workers

PostgreSQL

Persistent storage, indexing, transactions, query processing

Optional Redis Cache and Background Workers

Caching, queues, asynchronous tasks, rate-limit counters

In a typical deployment:

  1. The user sends an HTTPS request to the server.
  2. Nginx receives the request and routes it to the appropriate service.
  3. Next.js renders the frontend or serves a page, depending on the application's rendering strategy.
  4. FastAPI handles API requests, business logic, and database operations.
  5. PostgreSQL processes queries and returns the requested data.
  6. Redis, if configured, can provide cached results and reduce repeated database queries.
  7. Nginx delivers the response to the client.

For heavier workloads, these components can be distributed across multiple servers rather than running on a single Ubuntu machine.

2. Preparing Ubuntu Server for Heavy Loads

The server's operating system and hardware form the foundation of application performance. Even well-optimized application code can struggle if the server has insufficient memory, slow storage, or inadequate CPU capacity.

2.1 Choose appropriate server resources

The required resources depend on the application's workload, number of concurrent requests, database size, and processing requirements.

The following are illustrative starting configurations, not guaranteed capacity estimates.

WorkloadCPURAMStorage
Development or testing2 vCPU4 GB40 GB SSD
Small production4 vCPU8 GB80 GB NVMe
Medium production8 vCPU16–32 GB160 GB NVMe
High-load production16+ vCPU32–64+ GBHigh-performance NVMe

For database-intensive applications, memory and storage performance are particularly important. For CPU-intensive workloads, such as data processing or complex calculations, CPU capacity can become the primary bottleneck.

It is also important to leave sufficient memory for Ubuntu, Nginx, monitoring, and other system services.

2.2 Update and install essential packages

Start by updating Ubuntu and installing the basic tools required for server management.

sudo apt update
sudo apt upgrade -y

sudo apt install -y \
    nginx \
    postgresql \
    postgresql-contrib \
    python3 \
    python3-venv \
    python3-pip \
    nodejs \
    npm \
    git \
    curl \
    ufw \
    htop \
    sysstat \
    unzip

Install supported, compatible versions of Node.js and Python for the application's dependencies rather than assuming the Ubuntu repository version is appropriate for production.

2.3 Configure the firewall

Only expose the network ports required for public access and administration.

sudo ufw default deny incoming
sudo ufw default allow outgoing

sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp

sudo ufw enable
sudo ufw status

PostgreSQL, FastAPI, and Next.js should normally listen on localhost or a private network rather than being directly accessible from the public internet.

2.4 Configure system monitoring

Install and use monitoring utilities to understand server resource consumption.

htop

For memory:

free -h

For disk utilization:

df -h

For disk I/O:

iostat -xz 1

For system load:

uptime

Monitoring these metrics helps identify whether the application is CPU-bound, memory-constrained, or waiting on disk operations.

3. Optimizing FastAPI for High Concurrency

FastAPI is an asynchronous Python web framework designed for building APIs. It supports asynchronous request handling, dependency injection, data validation, and automatic API documentation.

However, FastAPI's asynchronous capabilities alone do not guarantee high throughput. The application must use appropriate concurrency, database connections, and resource management.

3.1 Use Uvicorn with multiple workers

Uvicorn is an ASGI server that runs FastAPI applications. A single Uvicorn process can handle multiple concurrent asynchronous requests, but CPU-bound workloads and process-level isolation may benefit from multiple workers.

For production, run FastAPI using multiple worker processes, managed by a process manager such as systemd or another supported deployment mechanism.

Install the required packages:

pip install fastapi uvicorn[standard]

An example command using Uvicorn's worker support through Gunicorn is:

gunicorn main:app \
    --worker-class uvicorn_worker.UvicornWorker \
    --workers 4 \
    --bind 127.0.0.1:8000 \
    --timeout 60 \
    --access-logfile -

This example assumes the uvicorn-worker package is installed:

pip install gunicorn uvicorn-worker

The number of workers should be based on CPU capacity, memory consumption, and load-testing results. Four workers may be suitable for a particular four-vCPU application, but it is not a universal optimal setting.

Each worker is a separate process and consumes its own memory and database connections.

3.2 Use asynchronous endpoints appropriately

FastAPI supports both synchronous and asynchronous endpoints.

An asynchronous endpoint is particularly useful when the request spends time waiting for external services, databases, or other I/O operations.

Example:

from fastapi import FastAPI
import httpx

app = FastAPI()

@app.get("/api/data")
async def get_data():
    async with httpx.AsyncClient() as client:
        response = await client.get(
            "https://example.com/data",
            timeout=10
        )

    return response.json()

For production applications, avoid creating a new HTTP client for every request. Instead, create and reuse a client through FastAPI's lifespan mechanism, allowing connection pooling and reducing repeated connection setup.

Use async def when the operations inside it are genuinely asynchronous. Calling blocking functions directly inside an asynchronous endpoint can block the event loop and reduce throughput.

For CPU-intensive operations, such as image processing, large calculations, or machine learning inference, consider a process pool or dedicated background workers.

3.3 Use connection pooling for database access

Opening a new PostgreSQL connection for every API request is inefficient and can exhaust database resources.

Connection pooling allows multiple requests to reuse a limited number of database connections.

For example, with SQLAlchemy's asynchronous engine:

from sqlalchemy.ext.asyncio import (
    create_async_engine,
    async_sessionmaker
)

DATABASE_URL = (
    "postgresql+asyncpg://appuser:password"
    "@127.0.0.1:5432/appdb"
)

engine = create_async_engine(
    DATABASE_URL,
    pool_size=10,
    max_overflow=5,
    pool_timeout=30,
    pool_recycle=1800,
    pool_pre_ping=True
)

SessionLocal = async_sessionmaker(
    engine,
    expire_on_commit=False
)

The database session can then be managed through a FastAPI dependency:

from fastapi import Depends
from sqlalchemy.ext.asyncio import AsyncSession

async def get_db():
    async with SessionLocal() as session:
        yield session

And used in an endpoint:

@app.get("/api/customers")
async def get_customers(
    db: AsyncSession = Depends(get_db)
):
    result = await db.execute(
        select(Customer).limit(100)
    )

    return result.scalars().all()

The example assumes that select, Customer, and SessionLocal are defined in the application.

The pool settings must be coordinated with the total number of application workers and the database's maximum connection capacity.

For example, four workers with a pool size of 10 and a maximum overflow of five could potentially request up to 60 connections in total. This needs to be considered alongside administrative connections and any other applications using PostgreSQL.

3.4 Avoid blocking operations

A common performance problem is executing synchronous, blocking operations inside asynchronous endpoints.

Examples include:

  • Using synchronous HTTP requests in async def endpoints.
  • Performing lengthy file operations on the event loop.
  • Running CPU-heavy calculations directly in request handlers.
  • Calling slow third-party services without timeouts.

Use asynchronous libraries for asynchronous workloads, or offload blocking operations to a thread pool or dedicated worker.

3.5 Implement pagination

Returning thousands of database records in a single API response increases query time, memory usage, network traffic, and frontend rendering time.

Instead, implement pagination.

@app.get("/api/customers")
async def get_customers(
    page: int = 1,
    page_size: int = 25,
    db: AsyncSession = Depends(get_db)
):
    page = max(page, 1)
    page_size = min(max(page_size, 1), 100)

    offset = (page - 1) * page_size

    result = await db.execute(
        select(Customer)
        .order_by(Customer.id)
        .offset(offset)
        .limit(page_size)
    )

    return result.scalars().all()

For very large tables, cursor-based pagination using indexed columns is often more efficient than increasingly large offsets.

3.6 Move lengthy operations to background workers

Tasks such as generating PDF reports, sending bulk emails, processing large files, or running lengthy data calculations should not keep HTTP requests open unnecessarily.

A typical architecture uses:

  • FastAPI to accept the request.
  • Redis or another message broker to queue the job.
  • A worker system such as Celery or RQ to process it.
  • PostgreSQL to store the job status and results, where appropriate.

The API can return a job identifier immediately, allowing the frontend to retrieve the status or results later.

FastAPI's built-in BackgroundTasks can be useful for lightweight tasks, but a durable external queue is generally more appropriate for long-running or critical work that must survive application restarts.

4. Optimizing Next.js for Performance

Next.js provides server-side rendering, static site generation, client-side navigation, and server components. How these features are used has a major impact on application performance.

4.1 Choose the right rendering strategy

Not every page needs to be rendered dynamically on every request.

Rendering strategySuitable use casesPerformance considerations
Static generationMarketing pages, documentation, articlesLow server processing for each visit
Incremental static regenerationFrequently updated contentRegenerates pages when required
Server-side renderingPersonalized dashboards, dynamic contentRequires server processing
Client-side renderingHighly interactive application viewsCan reduce initial server rendering, but depends on client resources

Static pages can be served directly by Nginx or a CDN, reducing the work performed by the Next.js server.

Dynamic rendering is useful for pages that depend on authentication, frequently changing information, or personalized content.

Choose the rendering approach based on actual page requirements rather than using server-side rendering for every page by default.

4.2 Optimize server components and data fetching

Next.js server components allow data to be fetched on the server without shipping the component's implementation to the browser.

For example:

export default async function CustomersPage() {
    const response = await fetch(
        "http://127.0.0.1:8000/api/customers",
        {
            cache: "no-store"
        }
    );

    if (!response.ok) {
        throw new Error("Unable to load customers");
    }

    const customers = await response.json();

    return (
        <main>
            <h1>Customers</h1>

            {customers.map((customer) => (
                <div key={customer.id}>
                    {customer.name}
                </div>
            ))}
        </main>
    );
}

This is an illustrative App Router example. The API URL and caching strategy should be adapted to the production architecture.

Avoid sequential requests when independent data can be retrieved concurrently:

const [customers, orders] = await Promise.all([
    getCustomers(),
    getOrders()
]);

This can reduce total page loading time when both requests are independent.

4.3 Reduce unnecessary client-side JavaScript

Large JavaScript bundles can increase initial page load time and browser CPU consumption.

To reduce unnecessary client-side processing:

  • Prefer server components when browser interaction is not required.
  • Use client components only for interactive functionality.
  • Dynamically load large components that are not immediately needed.
  • Avoid unnecessary dependencies.
  • Use optimized image formats and responsive image sizes.
  • Review bundle sizes during production builds.

For example, a complex charting component may be dynamically imported so it does not delay the initial rendering of a page where the chart is below the fold.

4.4 Optimize image delivery

Large images can consume considerable bandwidth and slow down page rendering.

Next.js provides an image optimization component that supports responsive image sizing and modern image formats, depending on the configured loader and deployment environment.

import Image from "next/image";

export default function Banner() {
    return (
        <Image
            src="/images/banner.webp"
            alt="Application dashboard"
            width={1200}
            height={500}
            priority
        />
    );
}

Use priority only for important above-the-fold images, such as a primary hero image, rather than applying it to every image.

For applications with many images, a CDN or dedicated image service may further reduce bandwidth and processing demands.

4.5 Run Next.js in production mode

Do not use the Next.js development server for production traffic.

Build the application:

npm run build

Run the production server:

npm run start

For a production deployment, manage the Next.js process through systemd or an appropriate process manager so that it restarts after unexpected failures and starts automatically after server reboots.

Where traffic is substantial, multiple Next.js instances can be run behind Nginx or a dedicated load balancer. Ensure that caching, session management, and any application state are compatible with multiple instances.

5. Configuring Nginx as a High-Performance Reverse Proxy

Nginx is a key component of the architecture. It handles incoming HTTP and HTTPS requests, routes traffic to the appropriate application, serves static content, and can distribute traffic across multiple application instances.

5.1 Use separate upstreams for Next.js and FastAPI

A typical application may run:

  • Next.js on 127.0.0.1:3000
  • FastAPI on 127.0.0.1:8000
  • Nginx on ports 80 and 443

Nginx can route API requests to FastAPI and all other requests to Next.js.

Example Nginx configuration for a single application instance

upstream nextjs_backend {
    server 127.0.0.1:3000;
    keepalive 32;
}

upstream fastapi_backend {
    server 127.0.0.1:8000;
    keepalive 32;
}

server {
    listen 80;
    server_name example.com www.example.com;

    return 301 https://example.com$request_uri;
}

server {
    listen 443 ssl;
    http2 on;

    server_name example.com;

    ssl_certificate
        /etc/letsencrypt/live/example.com/fullchain.pem;
    ssl_certificate_key
        /etc/letsencrypt/live/example.com/privkey.pem;

    client_max_body_size 20m;

    keepalive_timeout 30;
    sendfile on;
    tcp_nopush on;

    gzip on;
    gzip_vary on;
    gzip_min_length 1024;
    gzip_types
        application/json
        application/javascript
        application/xml
        text/css
        text/plain
        image/svg+xml;

    location /api/ {
        proxy_pass http://fastapi_backend;

        proxy_http_version 1.1;
        proxy_set_header Connection "";

        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For
            $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;

        proxy_connect_timeout 5s;
        proxy_send_timeout 60s;
        proxy_read_timeout 60s;
    }

    location / {
        proxy_pass http://nextjs_backend;

        proxy_http_version 1.1;
        proxy_set_header Connection "";

        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For
            $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;

        proxy_connect_timeout 5s;
        proxy_send_timeout 60s;
        proxy_read_timeout 60s;
    }
}

This configuration assumes that the application routes API requests under /api/ and that FastAPI's routes match that prefix. The trailing slash in proxy_pass and the API route definitions should be adjusted together if URL prefixes need to be rewritten.

For applications that use WebSockets, add the appropriate Upgrade and Connection headers in a dedicated location block.

5.2 Enable connection reuse

Nginx upstream keepalive connections reduce repeated TCP connection establishment between Nginx and the application servers.

The keepalive directive in the upstream configuration allows Nginx to retain idle upstream connections. The proxy configuration must also permit reuse, as shown above.

Connection reuse is especially useful when Nginx handles a large number of short API requests.

5.3 Configure worker processes

Nginx can handle many concurrent connections using its event-driven architecture.

A common baseline in /etc/nginx/nginx.conf is:

worker_processes auto;

events {
    worker_connections 4096;
    multi_accept on;
}

http {
    sendfile on;
    tcp_nopush on;
    tcp_nodelay on;

    keepalive_timeout 30;

    include /etc/nginx/mime.types;
    default_type application/octet-stream;

    include /etc/nginx/conf.d/*.conf;
    include /etc/nginx/sites-enabled/*;
}

worker_processes auto generally starts a worker for each available CPU core. worker_connections defines the maximum simultaneous connections per worker, subject to operating-system file descriptor limits and other constraints.

These settings do not represent the number of users the application can support. The actual capacity also depends on upstream connections, memory, application performance, and request duration.

Check the configuration before reloading:

sudo nginx -t
sudo systemctl reload nginx

5.4 Enable compression carefully

Gzip can reduce the size of text-based responses, including HTML, CSS, JavaScript, and JSON.

However, compressing already-compressed content, such as JPEG, WebP, or compressed video, generally provides little benefit and may waste CPU resources.

For very high traffic, evaluate whether compression should be handled by Nginx, a CDN, or both, and avoid redundant compression.

5.5 Implement rate limiting

Rate limiting can protect APIs from excessive requests, automated abuse, and accidental traffic spikes.

For example, define a rate-limit zone in the http block:

limit_req_zone $binary_remote_addr
    zone=api_limit:10m
    rate=10r/s;

Apply it to a specific API location:

location /api/ {
    limit_req zone=api_limit burst=20 nodelay;

    proxy_pass http://fastapi_backend;

    proxy_set_header Host $host;
    proxy_set_header X-Real-IP $remote_addr;
    proxy_set_header X-Forwarded-For
        $proxy_add_x_forwarded_for;
    proxy_set_header X-Forwarded-Proto $scheme;
}

The example limits requests based on client IP, but real-world rate limits should reflect endpoint sensitivity, legitimate traffic patterns, and authenticated user identity where applicable.

If Nginx is behind a CDN or another proxy, configure trusted client IP handling correctly. Otherwise, multiple users may share a proxy IP, or clients may be able to spoof their addresses.

6. Optimizing PostgreSQL for Large Databases

PostgreSQL is often one of the most important performance components of a data-intensive application. Poorly designed queries, missing indexes, excessive connections, and inefficient transactions can cause application-wide slowdowns.

6.1 Configure PostgreSQL memory

PostgreSQL uses memory for caching, query operations, and maintaining database connections.

Important settings include:

SettingPurpose
shared_buffersMemory used for PostgreSQL's shared data cache
effective_cache_sizePlanner estimate of available cache
work_memMemory available to individual query operations
maintenance_work_memMemory used for maintenance tasks
max_connectionsMaximum permitted database connections

An illustrative starting configuration for a dedicated PostgreSQL server with 16 GB RAM might look like:

shared_buffers = 4GB
effective_cache_size = 12GB
work_mem = 8MB
maintenance_work_mem = 512MB
max_connections = 100

These are not universal recommended values. For a shared application server, PostgreSQL may need substantially less memory to leave resources available for FastAPI, Next.js, and Nginx.

In particular, work_mem can be consumed by multiple query operations per connection, so setting it too high can result in excessive memory usage during concurrent workloads.

After modifying PostgreSQL settings, restart the service when required:

sudo systemctl restart postgresql

6.2 Create appropriate database indexes

Indexes allow PostgreSQL to locate rows without scanning entire tables.

Consider a table containing millions of customer records:

CREATE TABLE customers (
    id BIGSERIAL PRIMARY KEY,
    first_name VARCHAR(100),
    last_name VARCHAR(100),
    email VARCHAR(255),
    created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);

If the application frequently searches customers by email, create an index:

CREATE UNIQUE INDEX idx_customers_email
ON customers(email);

For queries that filter records by date:

CREATE INDEX idx_customers_created_at
ON customers(created_at);

For queries that filter by status and sort by date, a composite index may be appropriate:

CREATE INDEX idx_orders_status_created
ON orders(status, created_at DESC);

Indexes should reflect actual query patterns. Too many indexes increase storage usage and slow down inserts, updates, and deletes.

Use EXPLAIN ANALYZE to inspect query execution plans:

EXPLAIN ANALYZE
SELECT *
FROM customers
WHERE email = 'customer@example.com';

Review execution time, rows scanned, index usage, and estimated versus actual row counts.

6.3 Avoid SELECT *

Selecting only the columns required by the application reduces data transfer, memory usage, and unnecessary processing.

Instead of:

SELECT *
FROM customers;

Use:

SELECT id, first_name, last_name, email
FROM customers
LIMIT 100;

For large tables, combine selective queries with pagination and suitable indexes.

6.4 Use connection pooling

PostgreSQL is not designed to create unlimited simultaneous client connections.

An external connection pooler, such as PgBouncer, can help manage many application connections using a smaller number of database server connections.

Connection pooling architecture

FastAPI workers

Many concurrent API requests

PgBouncer

Queues and reuses database connections

PostgreSQL

Controlled number of database connections

PgBouncer supports different pooling modes, including session, transaction, and statement pooling. Transaction pooling can improve connection reuse, but some session-dependent features are incompatible with it. Check application requirements before selecting a mode.

6.5 Monitor slow queries

Enable PostgreSQL's slow-query logging where appropriate:

log_min_duration_statement = 500

This example logs statements taking at least 500 milliseconds. The threshold should be chosen according to the application's performance targets and expected workload.

PostgreSQL's pg_stat_statements extension can help identify queries that consume the most total execution time or are executed frequently.

For example:

SELECT
    query,
    calls,
    total_exec_time,
    mean_exec_time
FROM pg_stat_statements
ORDER BY total_exec_time DESC
LIMIT 20;

This helps prioritize query optimization based on actual database workload rather than assumptions.

6.6 Use appropriate transactions

Transactions help maintain data consistency, but long-running transactions can hold locks, delay vacuum operations, and contribute to contention.

Keep transactions short, avoid waiting for external HTTP requests while holding database transactions open, and use suitable isolation levels for the application's requirements.

Regularly monitor deadlocks, lock waits, table bloat, and autovacuum activity as the database grows.

7. Implementing Redis for Caching

Redis can reduce repeated database queries and improve response times for frequently accessed data.

Without caching, every request for the same information may trigger another database query. With caching, the application can retrieve a previously computed result from memory.

7.1 Typical caching architecture

FastAPI request

Redis cache

Is the requested data available?

Cache hit

Return cached data

Cache miss

Query PostgreSQL

PostgreSQL

Retrieve data and populate cache for future requests

7.2 Install Redis

On Ubuntu, Redis can be installed using:

sudo apt update
sudo apt install redis-server

Check whether Redis is running:

sudo systemctl status redis-server

Test the connection:

redis-cli ping

A successful connection returns:

PONG

Redis should be configured to listen only on localhost or a secured private network, with authentication and other appropriate protections where needed.

7.3 Implement caching in FastAPI

Install the Redis Python client:

pip install redis

Example using the asynchronous Redis client:

import json
import redis.asyncio as redis

from fastapi import FastAPI

app = FastAPI()

cache = redis.Redis(
    host="127.0.0.1",
    port=6379,
    decode_responses=True
)

@app.get("/api/dashboard")
async def get_dashboard():
    cached = await cache.get("dashboard")

    if cached:
        return json.loads(cached)

    data = {
        "total_customers": 2500,
        "total_orders": 8500
    }

    await cache.set(
        "dashboard",
        json.dumps(data),
        ex=300
    )

    return data

This simplified example uses fixed data to demonstrate the caching mechanism. In a real application, the cache miss branch would retrieve the data from PostgreSQL.

The ex=300 parameter sets a five-minute time-to-live, after which the cached value expires.

7.4 Choose what to cache

Caching is particularly useful for:

  • Frequently accessed dashboard statistics.
  • Product catalogs and category lists.
  • Public content and articles.
  • Application configuration.
  • Expensive aggregation queries.
  • Rate-limiting counters.

Not all data should be cached. Financial transactions, sensitive user-specific data, and frequently changing records require careful cache design.

Cache invalidation is also important. When data changes, the application should either invalidate the relevant cache entry or update it. Otherwise, users may receive outdated information.

For personalized responses, include the appropriate user or tenant identifier in the cache key to prevent data from being shared across users.

8. Using Background Processing and Message Queues

Some workloads do not need to be completed while the user is waiting for an HTTP response.

Examples include:

  • Sending notification emails.
  • Generating invoices and reports.
  • Processing uploaded documents.
  • Importing large CSV files.
  • Synchronizing third-party APIs.
  • Running data analytics.

Instead of processing these tasks directly inside an API request, place them in a background queue.

Next.js frontend

FastAPI

Accept request and enqueue job

Redis / Message Queue

Store pending jobs

Worker 1

Worker 2

Workers process jobs independently and store their results or status.

Tools such as Celery and RQ can manage background tasks using Redis or compatible brokers.

For reliable processing, design jobs to be retryable and idempotent where possible. Track job states, failures, retry counts, and completion times. This prevents temporary service interruptions from resulting in lost or duplicated business operations.

9. Scaling the Application Horizontally

When one server can no longer meet the application's performance targets, adding more resources to the same server may help, but eventually horizontal scaling becomes necessary.

Horizontal scaling means running multiple instances of an application and distributing incoming requests among them.

9.1 Single-server deployment

Internet traffic

Ubuntu Server

Nginx

Next.js

FastAPI

PostgreSQL

Redis

A single-server architecture is relatively straightforward to deploy and maintain. It can be sufficient for many small and medium applications, depending on actual workload.

However, it introduces a single point of failure. If the server goes offline, the entire application may become unavailable.

9.2 Multi-server deployment

Internet traffic

Load Balancer / Nginx

TLS termination, routing, health checks

Application Server 1

Next.js

FastAPI

Application Server 2

Next.js

FastAPI

PostgreSQL Server

Redis Server

A multi-server deployment separates application processing from database storage and other shared services.

This allows application instances to scale independently, while centralized data services maintain shared application state.

Important considerations include:

  • Use health checks to avoid routing requests to unavailable application instances.
  • Keep user sessions and application state outside individual instances, or use a compatible shared session strategy.
  • Configure database connection pools across all application servers.
  • Use private networking and access controls between servers.
  • Plan for PostgreSQL high availability and backup recovery.
  • Distribute background jobs across separate worker instances if needed.

Nginx can distribute requests among multiple upstream servers using round-robin, least-connections, or other supported methods. For larger deployments, a managed load balancer or dedicated load-balancing tier may be appropriate.

9.3 Scale PostgreSQL independently

Database scaling requires particular attention because it contains shared application state.

Potential approaches include:

  • Increasing CPU, RAM, and storage performance on the primary database server.
  • Optimizing queries and indexes before increasing hardware.
  • Using PgBouncer to manage connections.
  • Using read replicas for suitable read-heavy workloads.
  • Partitioning very large tables when the access patterns justify it.
  • Using an appropriate high-availability solution with tested failover procedures.

Read replicas can help distribute read queries, but replication lag means they may not immediately reflect recent writes. Transactions and read-after-write operations may need to be routed to the primary database.

10. Securing the Application

Security is essential for a high-load production application. A compromised server or resource-exhaustion attack can cause outages regardless of the application's performance optimizations.

10.1 Secure application services

  • Bind FastAPI and Next.js to localhost or private network interfaces.
  • Restrict PostgreSQL access to trusted application hosts.
  • Use SSH keys and disable password-based SSH authentication where operationally appropriate.
  • Use least-privilege Linux and PostgreSQL accounts.
  • Store secrets in protected environment files or a secrets management system.
  • Apply operating-system and dependency security updates.
  • Use TLS for public connections and secure internal communication where appropriate.

10.2 Protect API endpoints

FastAPI applications should implement appropriate authentication and authorization, input validation, request-size limits, and protection against abusive traffic.

Consider:

  • OAuth 2.0 or other suitable authentication mechanisms.
  • Role-based or permission-based access control.
  • Secure password hashing.
  • Short-lived access tokens where applicable.
  • Refresh-token rotation and revocation where applicable.
  • Request validation and safe database query construction.
  • Per-user and per-IP rate limiting.

Do not rely exclusively on frontend checks to enforce access control. All sensitive permissions must be validated by the backend.

10.3 Protect PostgreSQL

Create a dedicated database user for the application instead of using the PostgreSQL superuser.

For example:

CREATE USER appuser WITH PASSWORD 'replace-with-a-strong-secret';

GRANT CONNECT ON DATABASE appdb TO appuser;

Connect to appdb as a database administrator, then grant only the necessary schema and table privileges:

GRANT USAGE ON SCHEMA public TO appuser;

GRANT SELECT, INSERT, UPDATE, DELETE
ON ALL TABLES IN SCHEMA public
TO appuser;

GRANT USAGE, SELECT
ON ALL SEQUENCES IN SCHEMA public
TO appuser;

For future tables and sequences, configure default privileges for the role that creates them:

ALTER DEFAULT PRIVILEGES
FOR ROLE dbowner
IN SCHEMA public
GRANT SELECT, INSERT, UPDATE, DELETE
ON TABLES TO appuser;

ALTER DEFAULT PRIVILEGES
FOR ROLE dbowner
IN SCHEMA public
GRANT USAGE, SELECT
ON SEQUENCES TO appuser;

Replace dbowner with the actual role used to create application objects. Restrict access further for applications that only need read or limited write permissions.

10.4 Add protection against excessive traffic

Rate limiting at Nginx, application-level quotas, and suitable CDN or web application firewall protections can help control abusive requests.

For large public applications, a CDN may also provide distributed traffic absorption, caching, and additional security features.

Rate limiting should not replace proper capacity planning or infrastructure monitoring.

11. Monitoring and Performance Observability

A production application needs continuous monitoring to identify performance bottlenecks before they cause significant service interruptions.

Monitoring should cover the operating system, Nginx, FastAPI, Next.js, PostgreSQL, and Redis.

11.1 Essential monitoring tools

ToolPurpose
PrometheusCollects time-series metrics
GrafanaVisualizes metrics and dashboards
NetdataReal-time server monitoring
LokiCentralized log collection
SentryTracks application errors
pg_stat_statementsIdentifies expensive database queries
Implementing Centralized Monitoring and Alerting for Multiple Servers using Prometheus + Grafana | by Juwono (ジュウオノ) | Medium
Prometheus All Metrics | Grafana Labs
How to deploy the Netdata network and server monitor on Linux

11.2 Monitor key performance indicators

Recommended metrics

MetricWhat it indicates
CPU utilizationApplication processing pressure and CPU saturation
Memory utilizationMemory pressure, cache usage, and potential out-of-memory conditions
Disk I/O latencyStorage bottlenecks and database I/O performance
API response timeHow quickly API requests are processed
Requests per secondApplication throughput
HTTP error rateApplication failures and upstream issues
Database connectionsConnection usage and potential exhaustion
Cache hit ratioHow effectively cached data is being reused
Queue depthWhether background workers are keeping up with incoming jobs

Monitor latency percentiles, particularly p95 and p99, rather than relying only on average response time. A low average can conceal slow responses affecting a significant portion of users.

Define service-level objectives (SLOs) for response times, availability, and error rates based on the application's actual requirements.

12. Load Testing Before Going Live

Load testing helps establish how much traffic the application can handle and where performance begins to degrade.

Testing should be performed in an isolated staging environment or with controlled production testing that will not disrupt real users.

12.1 Use load-testing tools

Popular tools include:

  • k6 — scriptable performance and load testing.
  • Locust — Python-based user simulation.
  • Apache JMeter — configurable performance and protocol testing.

12.2 Example using k6

A simple load-testing script can simulate concurrent requests to a FastAPI endpoint.

import http from "k6/http";
import { check, sleep } from "k6";

export const options = {
    stages: [
        { duration: "1m", target: 50 },
        { duration: "3m", target: 100 },
        { duration: "3m", target: 250 },
        { duration: "1m", target: 0 }
    ],
    thresholds: {
        http_req_failed: ["rate<0.01"],
        http_req_duration: ["p(95)<500"]
    }
};

export default function () {
    const response = http.get(
        "https://example.com/api/customers"
    );

    check(response, {
        "status is 200": (r) => r.status === 200
    });

    sleep(1);
}

Run the test with:

k6 run load-test.js

This example ramps up to 250 virtual users. Virtual users are not the same as requests per second or actual simultaneous server-side requests, so interpret the results in the context of the script's request rate and response times.

The sample thresholds specify a target of less than 1% failed requests and a p95 response time below 500 ms. They are illustrative targets, not requirements for every application.

12.3 Test different traffic patterns

Suggested load-testing stages

  1. Baseline testEstablish normal response times and resource consumption with a small number of users.
  2. Load testGradually increase traffic to determine how the application behaves under expected workloads.
  3. Stress testIncrease traffic beyond expected capacity in a controlled environment to identify bottlenecks and failure behavior.
  4. Spike testSimulate sudden increases in traffic, such as a product launch or a large promotional campaign.
  5. Soak testRun a sustained workload over a longer period to identify memory leaks, connection exhaustion, and resource degradation.

During every test, monitor application response times, CPU utilization, memory, PostgreSQL query latency, database connections, and Nginx error logs.

The purpose is not simply to reach a large number of users, but to understand the conditions under which the application continues to meet its performance and reliability requirements.

13. Deploying Services with systemd

On an Ubuntu server, systemd can manage Next.js and FastAPI as persistent services. This ensures they start on reboot and can be restarted when they fail.

13.1 Create a FastAPI service

Create a systemd unit:

sudo nano /etc/systemd/system/fastapi.service

Example:

[Unit]
Description=FastAPI Application
After=network.target postgresql.service

[Service]
User=www-data
Group=www-data

WorkingDirectory=/var/www/myapp/backend

EnvironmentFile=/etc/myapp/backend.env

ExecStart=/var/www/myapp/backend/venv/bin/gunicorn \
    main:app \
    --worker-class uvicorn_worker.UvicornWorker \
    --workers 4 \
    --bind 127.0.0.1:8000 \
    --timeout 60

Restart=always
RestartSec=5

NoNewPrivileges=true
PrivateTmp=true

[Install]
WantedBy=multi-user.target

Ensure that www-data has the necessary read and execute permissions for the application directory and appropriate access to its environment file. Adjust the worker count based on testing.

Activate the service:

sudo systemctl daemon-reload

sudo systemctl enable fastapi
sudo systemctl start fastapi

sudo systemctl status fastapi

13.2 Create a Next.js service

Create another service:

sudo nano /etc/systemd/system/nextjs.service

Example:

[Unit]
Description=Next.js Application
After=network.target

[Service]
Type=simple

User=www-data
Group=www-data

WorkingDirectory=/var/www/myapp/frontend

Environment=NODE_ENV=production
Environment=PORT=3000
EnvironmentFile=/etc/myapp/frontend.env

ExecStart=/usr/bin/npm run start

Restart=always
RestartSec=5

NoNewPrivileges=true
PrivateTmp=true

[Install]
WantedBy=multi-user.target

Confirm the correct path to the Node.js and npm executables using:

which node
which npm

Then enable and start the service:

sudo systemctl daemon-reload

sudo systemctl enable nextjs
sudo systemctl start nextjs

sudo systemctl status nextjs

For production, use a consistent Node.js installation path and ensure the application has been built before starting the service.

13.3 Review application logs

For FastAPI:

sudo journalctl -u fastapi -f

For Next.js:

sudo journalctl -u nextjs -f

For Nginx:

sudo tail -f /var/log/nginx/access.log
sudo tail -f /var/log/nginx/error.log

For PostgreSQL:

sudo journalctl -u postgresql -f

Centralized log collection becomes increasingly useful when the application runs across multiple servers.

14. Backups and Disaster Recovery

Performance and scalability are only part of a reliable production deployment. Data protection and recovery planning are equally important.

14.1 Automate PostgreSQL backups

A basic PostgreSQL logical backup can be created using:

pg_dump -U appuser -d appdb \
    -F c \
    -f /backup/appdb.dump

Restore it using:

pg_restore \
    -U appuser \
    -d appdb \
    /backup/appdb.dump

For larger databases or more stringent recovery requirements, consider physical backups and continuous archiving with PostgreSQL's write-ahead logs (WAL).

14.2 Follow the 3-2-1 backup principle

Maintain:

  • Three copies of important data.
  • Two different types of storage.
  • One copy stored off-site.

Backups should be encrypted where appropriate, access-controlled, and protected from accidental deletion or ransomware.

14.3 Test restoration

A backup is useful only if it can be restored successfully.

Periodically restore backups to an isolated environment and verify database consistency, application functionality, and recovery time.

Define:

  • RPO (Recovery Point Objective): The maximum acceptable amount of data loss measured in time.
  • RTO (Recovery Time Objective): The target time to restore the service after a disruption.

These requirements should influence the backup frequency, replication strategy, and disaster recovery architecture.

15. Recommended Production Architecture

For an application that is expected to experience substantial traffic, the following is an example of a distributed deployment.

Users and Internet Traffic

Web browsers, mobile apps, external API clients

CDN / Load Balancer

TLS, caching, traffic distribution, security controls

App Server 1

Nginx

Next.js

FastAPI

App Server 2

Nginx

Next.js

FastAPI

PostgreSQL

Primary database and optional read replicas

Redis

Caching and queue infrastructure

Background Workers

Independent processing for long-running jobs

Centralized Monitoring and Logging

Metrics, alerts, tracing, logs, and operational dashboards

This architecture allows different parts of the application to scale independently. For example, more application servers can be added when API traffic increases, while background worker capacity can be increased separately when processing queues grow.

The database should be scaled based on query workload and data access patterns rather than simply matching the number of application servers.

16. Practical Performance Optimization Checklist

Production readiness

0 of 14 completeUse production builds for Next.js and a production ASGI server for FastAPIConfigure Nginx as a reverse proxy with HTTPS and connection reuseUse asynchronous I/O for suitable FastAPI endpointsImplement database connection pooling and sensible connection limitsAdd indexes based on real query patterns and inspect slow queriesUse pagination and limit unnecessary data retrievalApply suitable Next.js rendering strategies and reduce client-side JavaScriptIntroduce Redis caching for frequently accessed data where appropriateMove lengthy tasks to background workersConfigure rate limiting and secure all public-facing endpointsMonitor CPU, memory, disk, database, and application metricsRun load, stress, spike, and soak tests in a controlled environmentSet up automated backups and regularly test restorationEstablish performance objectives, alerting, and incident response proceduresReset checklist

Conclusion

Building a high-performance FastAPI, Next.js, PostgreSQL, and Nginx application on Ubuntu requires a combination of efficient application design, optimized database operations, reliable infrastructure, and continuous monitoring.

FastAPI provides asynchronous API processing, Next.js supports multiple rendering strategies for fast web experiences, PostgreSQL offers robust relational data management, and Nginx provides efficient request routing and traffic handling.

Redis caching, database connection pooling, background processing, and horizontal scaling can further improve performance when they address measured bottlenecks.

The most important principle is to optimize based on actual workload measurements rather than assumptions. Begin with a well-configured single-server deployment when appropriate, establish performance benchmarks, identify bottlenecks, and scale individual components as the application's traffic and operational requirements grow.

With suitable architecture, testing, monitoring, and recovery procedures, this technology stack can support demanding production workloads while remaining maintainable and adaptable as the application expands.





Popular Categories

Agile 2 Android 2 Artificial Intelligence 52 Backup Tools 2 Blockchain 2 Cloud Storage 5 Code Editors 2 Computer Languages 12 Cybersecurity 9 Data Science 18 Database 9 Digital Marketing 3 Ecommerce 4 Email Server 2 FastAPI 1 Finance 3 Google 6 HTML-CSS 2 Industries 6 Infrastructure 5 iOS 3 IoT 1 Javascript 6 Latest Technologies 46 Linux 10 LLMs 11 Machine Learning 32 Mobile 3 Msths & Stats 1 MySQL 3 Open source 1 Operating Systems 7 PHP 2 PostgreSQL 1 Project Management 3 Python Programming 28 SEO - AEO 5 Software Development 50 Software Testing 3 Web Server 8 Work Ethics 2
Recent Articles
Build a High-Performance FastAPI, Next.js, PostgreSQL, and Nginx Application for Heavy Loads on Ubuntu Server
FastAPI

The Chief AI Officer Role
Artificial Intelligence

PACS vs. CAMT: Decoding the Core Messaging Categories of ISO 20022
Data Science

UFW vs. Firewalld: Firewall Comparison
Cybersecurity

Ubuntu vs. AlmaLinux: Server OS Choice
Linux

A Production-Ready Backup Architecture
Backup Tools

rsync vs rclone vs restic — A Practical, In-Depth Comparison
Backup Tools

A Comprehensive Guide to rsync
Cloud Storage

Tabulator Data Grid
Data Science

ESP32: Capabilities and Practical Project Examples
IoT

Prompt Engineering: Chain-of-Thought vs Tree-of-Thought
Artificial Intelligence

Prompt Engineering: Zero-Shot vs Few-Shot Prompting
Data Science

Lambda Functions in Python Programming
Python Programming

functools in Python Programming
Latest Technologies

MySQL Database Sharding: A Comprehensive Guide to Horizontal Scaling
Database

Database Sharding: Scaling Horizontally for Modern Applications
Database

Best Python Packages to Learn in 2026
Artificial Intelligence

Step-by-Step Guide to Google Play Store Submission
Google

Step-by-Step Guide to App Store Submission
iOS

Google Nano Banana: The AI Image Tool That Took the Internet by Storm
Artificial Intelligence

Best Practices For Software Development Using Google Gemini 2.5 Pro Through Prompt Engineering
Data Science

Email-Based Passcode Authentication: A Secure and User-Friendly Approach
Software Development

AI Hot Topics Mid-2025
Artificial Intelligence

The Top 3 Python Web Frameworks for 2025: Django, FastAPI, and Flask
Python Programming

Best NLP Libraries for Natural Language Processing in 2025
Artificial Intelligence

Python Implementation of a Simple Blockchain
Blockchain

Explain blockchain like I’m a 10-year-old, using simple analogies.
Blockchain

Prompt Engineering: The Art of Communicating with AI
Artificial Intelligence

Best Generative AI Tools for Code Generation
Artificial Intelligence

TensorFlow vs PyTorch: A Comprehensive Comparison
Artificial Intelligence