Documentation

What the platform does, how to run it, and why each decision was made the way it was. Written to be read in order, and to be usable out of it.

Overview

What this is, what shape it has, and what it deliberately does not do.

What it is

DCManage is a datacenter and network management platform: it holds the estate — sites, racks, switches, ports, servers and the services running on them — measures traffic against it, and records who changed what.

It is a Laravel application backed by PostgreSQL. It is installed on a machine you control, on your own network, and it reaches nothing on the public internet unless you configure it to.

The shape of it

Three processes, deliberately separated:

  • The web tier serves the interface. It cannot open a stored device credential, which means no page can be made to wait on a device that has stopped answering.
  • The scheduler runs collection and maintenance on a timetable, each task in its own process with its own overlap guard.
  • The worker performs device work. Anything that touches a switch is queued to it and reported back.

That separation is not architecture for its own sake. A monitoring host that became slow once held every web worker on a panel until the whole thing stopped answering; this arrangement makes that impossible rather than unlikely.

What it does not do

It does not issue invoices. It measures what a customer used and on what basis they are billed; the billing system remains whatever you already have.

It does not require an agent on a customer machine. Everything is read from network equipment.

It does not act on a customer unless you switch enforcement on, and it ships with that switched off.

Installation

What it needs, the two identities it runs as, and how a release is applied.

Requirements

  • Linux, PHP 8.3 or later with the SNMP extension
  • PostgreSQL 16 or later — the schema uses declarative partitioning and privilege revocation that are not portable, and are not meant to be
  • Network reachability to the equipment being read, directly or through a proxy

Two identities

The application runs as two accounts with different reach, and the difference is the point.

  • www-data serves HTTP and cannot read the master key.
  • A worker account runs the scheduler and the queue and can.

The master key opens every stored device credential. Because the web tier cannot read it, a compromised page cannot produce a switch password, and no request can be made to hold a socket open to a device.

A page that says it cannot open the stored credential is not a fault. It is a caller that has not been routed through the queue.

Applying a release

Files are synchronised to the server and an apply script does the rest: it restores ownership, migrates, clears caches, restarts the web process and the worker, then verifies every invariant and exits non-zero if one fails.

Restarting the worker is not optional. It is long-lived and holds the classes it loaded, so a deploy that leaves it running keeps executing the previous release.

How traffic is measured

What is read, what is recorded when a reading is missed, and how a figure becomes an invoice line.

Collection

The unit of work is the device, not the port. The collector builds one request from the interface indexes it already knows about rather than walking the interface table — on a 349-interface chassis that is the difference between two seconds and a seventh of one.

Counters are 64-bit where the device offers them. A counter that wraps or resets is detected and recorded as such rather than producing a spike.

Gaps and quality

Every sample carries a quality: usable, spanning a gap, following a reset, or suspect. A window that nobody watched reports nothing rather than zero, and coverage describes what was actually observed rather than what was inferred.

This matters at the moment somebody disputes a bill. A figure that cannot say how much of its window it actually saw is not evidence.

The billing basis

A service is billed on download, on upload, or on both. The setting is resolved from the service, falling back to the server and then the site, and the percentile is taken on that direction.

Taking the larger of the two directions is a common shortcut and it is wrong for anyone selling transit: it charges a customer for traffic they were never invoiced for.

External monitoring

Running an existing system alongside this one, and why the comparison is the point.

Why run both

A new measurement system asks to be trusted with invoices on its first day. It should not be. Registering the system you already have as a source lets the same port be read twice and the two figures compared over a full cycle.

When they agree, you have evidence. When they disagree, you have a question worth answering before it reaches a customer.

Adding a source

A source is an address, a credential and how to present it. Its sensors are cached locally so the interface never asks a remote system anything while somebody is waiting, and traffic sensors are attached to ports by matching what both systems agree on.

Matching is deliberately narrow. An unmatched sensor is left alone and counted, because a blank column says "we do not know" and a wrong one says something confident about the wrong customer.

When it stops answering

After repeated connection failures the platform stops asking. It does not then wait out a timer: it opens a bare socket — no encryption, no credentials, no API — and only issues a real request if that answers. Against a host with no route that costs about a millisecond.

An error that is not a connection failure does not count toward that. A system answering "not authorised" is reachable and has a credential problem, and hiding that behind a timeout sends somebody to look at the wrong thing.

Proxies

Reaching a range the platform has no route to, and what a proxy cannot carry.

When you need one

Management controllers usually live on a private range that the management server has no route to. No credential fixes that; a route is what is missing.

A proxy is registered once and named by the equipment reached through it. Equipment that names none is connected to directly, which is what everything on the platform's own subnet uses.

Kinds

  • SOCKS5 — the usual answer, with optional credentials and name resolution at the far end.
  • SOCKS4 — for older equipment; SOCKS4a is used automatically when a name rather than an address is being reached.
  • HTTP CONNECT — tunnelling through a web proxy.
  • GRE — a kernel tunnel to the jump host. Ranges on that record become routes. SNMP, ICMP and HTTPS all take it. SOCKS remains the fallback when the tunnel is down.

Resolving names at the proxy is usually required for SOCKS: a name that only exists on the private range is one the platform has no way to look up.

What it will not carry

SNMP does not go through SOCKS or HTTP CONNECT. It is UDP through a PHP extension that hands out no socket, so there is nothing to put a stream proxy underneath.

A GRE tunnel is a route, not a stream proxy, so SNMP, ICMP and HTTPS all take it. Keep SOCKS as well: when the tunnel is down, HTTPS still has a path and SNMP does not until it returns.

Standing one up

Any SOCKS5 server will do. These instructions use Dante, because it authenticates and — the part that matters — it can refuse a destination, which is what turns a hole in your network boundary into a door with a lock on it.

Put it on a machine that already reaches the management range: usually a jump host or the monitoring server, not the platform. If the platform could reach the range it would not need a proxy.

1. Install it, and make an account for the platform to use. The account needs no shell and no home directory — it exists to be a name and a password.

apt install dante-server
useradd --system --no-create-home --shell /usr/sbin/nologin dcmanage-proxy
passwd dcmanage-proxy

2. Write the rules. This is the whole of /etc/danted.conf. Replace eth1 with the interface facing the management range, and 185.81.96.142 with the platform’s own address.

logoutput: /var/log/sockd.log
internal: 0.0.0.0 port = 1080
external: eth1

# Username and password, checked against system accounts.
socksmethod: username
user.privileged: root
user.unprivileged: nobody

# Who may talk to this proxy at all: the platform, and nothing else.
client pass {
    from: 185.81.96.142/32 to: 0.0.0.0/0
    log: connect disconnect error
}
client block {
    from: 0.0.0.0/0 to: 0.0.0.0/0
    log: connect error
}

# Where it may go once it has. One rule per destination range — copy the
# block, change the CIDR and the port, leave the rest. A single rule to
# 0.0.0.0/0 is an open proxy.
socks pass {
    from: 185.81.96.142/32 to: 172.16.30.0/24 port = 443
    command: connect
    socksmethod: username
    log: connect disconnect error
}
socks pass {
    from: 185.81.96.142/32 to: 172.16.31.0/24 port = 443
    command: connect
    socksmethod: username
    log: connect disconnect error
}
socks pass {
    from: 185.81.96.142/32 to: 10.0.0.0/24 port = 22
    command: connect
    socksmethod: username
    log: connect disconnect error
}

# Everything else, refused and written down.
socks block {
    from: 0.0.0.0/0 to: 0.0.0.0/0
    log: connect error
}

3. Start it, and make it survive a reboot.

systemctl enable --now danted
systemctl status danted

Two things about the shipped unit will bite, and both are silent.

It sets ReadOnlyDirectories=/var, so the logoutput line above cannot open its file and the daemon runs with no log at all — which means the block rules refuse without writing anything down, and a refusal you cannot see is a refusal you will misread as a network fault. It also waits only for network.target, which is reached before an interface has an address; if internal names an address rather than 0.0.0.0, a cold boot fails to bind and the unit dies with the proxy down until somebody notices.

mkdir -p /etc/systemd/system/danted.service.d
cat > /etc/systemd/system/danted.service.d/override.conf <<'EOF'
[Unit]
After=network-online.target
Wants=network-online.target

[Service]
ReadWritePaths=/var/log/sockd.log
Restart=on-failure
RestartSec=5s
EOF
touch /var/log/sockd.log
systemctl daemon-reload && systemctl restart danted
tail /var/log/sockd.log

The last line is the test. An empty file after a restart means the log never opened.

On Ubuntu, apt may refuse to install anything at all while a file:///cdrom line is left in /etc/apt/sources.list. Comment it out.

Port 1080 must not be reachable from anywhere but the platform. The client pass rule is Dante’s half; a host firewall on 1080 is the other half. Both. The two block rules are not decoration. Dante refuses by default when no rule matches, but writing the refusal down puts it in the log — and a proxy that refuses silently is a proxy you will spend an afternoon debugging from the wrong end.

Declaring the ranges, twice

The same ranges go in two places, and that is deliberate.

In the proxy server, as the socks pass rules above. This is the real boundary: enforced by a different program on a different machine, and it holds even if this platform is compromised outright. Ports belong here — HTTPS to an iLO range is not SSH to a jump range.

In this platform, on the proxy’s own record, one CIDR to a line, matching those rules. This one is not security — it is what catches a mistyped address before a connection is opened, with a message naming the proxy and the ranges it was declared for, instead of a timeout somebody spends twenty minutes on.

On the proxy list the same file is printed next to the records, so the two copies can be compared without leaving the page. A proxy with no ranges declared here will reach anything it is pointed at, and the proxy list says which ones are in that state. Every proxy created before this feature existed is in it. That is not an error, but it is worth fixing.

One consequence worth knowing: a proxy with declared ranges refuses a name as well as an address outside them. A name resolves at the far end, so nothing here can tell what it will become — and knowing where the request lands is the entire point.

Checking it works

Test the proxy from the platform’s own machine, as the account that will use it, before adding it here. A proxy that works from your laptop and not from the platform is the commonest way to lose an hour.

# Should answer: a controller on the declared range.
curl -sk --socks5-hostname dcmanage-proxy:PASSWORD@proxy.example:1080 \
     https://172.16.30.119/redfish/v1/Systems/1 | head

# Should be refused by the proxy, not by the platform.
curl -sk --socks5-hostname dcmanage-proxy:PASSWORD@proxy.example:1080 \
     https://example.com/

If the first answers and the second does not, the rules are right. If both answer, the socks block rule is missing or a socks pass is wider than you meant.

Once it is added here, the proxy list has a Reach button. It opens a bare socket to the proxy and nothing else — it cannot test the credential, because the web tier cannot read the master key and so cannot open a stored password. That is a property of the design rather than a gap in the button, and it is why the curl above is worth doing.

Access and audit

Who may do what, and the record of what they did.

Roles and exceptions

Permissions are declared in one registry that drives three consumers: the authorisation gate, the interface, and the role editor. That is what keeps them from drifting — a rendered button is always a pressable button, because visibility and permission read the same key.

A role is a named set. Individual exceptions sit on the administrator on top of their roles, and a revocation always wins over a role that grants the same thing.

The audit log

Every action that changed something, and every refusal. Entries cannot be edited or removed by anyone, including a full administrator: the UPDATE and DELETE privileges are revoked at the database rather than merely absent from the code.

Values that could carry a secret are redacted before storage. An audit trail that leaks a switch password is a credential store with a search interface.

Where it can be reached from

The panel can be restricted to a list of source addresses, and a refused attempt is recorded like any other action. That list is enforced before authentication, so an address that is not on it never reaches the sign-in form.

Security

How credentials are stored, why the web tier cannot read them, and what stays switched off.

Stored credentials

Device credentials are sealed with XChaCha20-Poly1305 under a master key held outside the application directory. The table, row and field are part of what each value is sealed against, so a ciphertext copied from one row to another does not decrypt.

A credential cannot be displayed once stored. The interface reports that one is held, which is a different statement from showing it.

Separated by identity

The web tier cannot read the master key. That is enforced by file ownership and asserted on every deploy, and it is the reason a slow or hostile device cannot exhaust the web process pool: those processes have no way to open a connection to one.

Enforcement stays off

The platform ships observing. It works out what it would have decided and writes it down, and changes nothing. Arming it requires a full administrator and a typed confirmation phrase.

Two systems acting on the same devices from different pictures is how a paying customer goes offline, so the platform assumes it is the second one until told otherwise.