Postman Test Data: Generate Reproducible Data Files for Newman
Postman test data usually lives in one of two places. Either it's a
hand-written data file with three or four rows nobody wants to extend,
or it's spread across request bodies as {{$randomFirstName}} and
{{$randomEmail}}, producing new values on every run. The first covers
too little, and the second can't be replayed: when iteration 7 fails in
CI, the values that broke it are gone. This post builds a third option.
One template generates the whole data file from a fixed seed. Valid rows
and negative rows each carry the status code they should produce, and
the same file runs identically in the Collection Runner and in Newman.
JsonFabrica's part in this is narrow. It's a JSON-over-HTTP generation
API. You call it with curl, save the data field of the response as
data.json, and Postman takes it from there. There's no Postman
integration, no Newman plugin, and nothing to install in your
collection.
How a Postman data file drives the Collection Runner and Newman
A data file turns one collection into a table-driven test. The Collection Runner, and Newman on the command line, run the whole collection once per row and expose that row's values as variables for the duration of the iteration.
A Postman data file is either a CSV with a header row or a JSON file containing an array of objects:
[
{ "email": "[email protected]", "plan": "pro", "expectedStatus": 201 },
{ "email": "not-an-email", "plan": "pro", "expectedStatus": 422 }
]
Each object is one iteration. Inside that iteration:
- Request fields reference values as
{{email}}or{{plan}}in the URL, headers, or body, the same syntax as environment variables. - Scripts read them with
pm.iterationData.get('email').pm.info.iterationgives the current iteration index, starting at 0.
Prefer JSON over CSV for test cases. JSON is the better choice because
the types are in the file: 201 is a number and true is a boolean.
With CSV, types are guessed: Newman converts unquoted numbers but leaves
true as the string "true", so a boolean check needs a conversion
first.
Hand-written data files vs. Postman dynamic variables
Both common approaches to Postman test data run into limits as soon as the data needs to be reviewed or replayed.
Hand-written data files stay small. Typing JSON rows is tedious, so the file settles at three to five rows, usually all happy-path. Negative cases end up as separate requests with hard-coded bodies, or don't get written. Nobody adds the twentieth row.
Dynamic variables aren't reproducible. Postman ships dozens of
built-in dynamic variables, such as {{$guid}}, {{$timestamp}},
{{$randomInt}}, {{$randomFirstName}}, and {{$randomEmail}}. Every
reference resolves to a fresh value, and Postman doesn't document a way
to seed them. Two consequences follow:
- A failing run can't be replayed. Whatever
{{$randomEmail}}produced on the failing iteration was never written down, unless you logged it. - Values in one iteration can't depend on each other. Each reference is
drawn independently, so nothing ties
{{$randomInt}}seats to the plan you sent or to the status code you expect. Validation rules are usually about exactly those relationships: a free plan allows one seat, and a team plan allows up to 50.
Dynamic variables are still useful for some jobs, and the section on when to use them covers which. They aren't a substitute for a set of test cases you can read, review, and rerun.
Postman test data where every row carries its expected outcome
The fix is to treat the data file as a list of test cases. Every row says what to send and what should happen:
| Field | Purpose |
|---|---|
case |
A readable label, used in test names |
name, email, plan, seats |
The request payload |
expectedStatus |
201 for valid rows, 422 for invalid ones |
expectedField |
For invalid rows, the field the API should flag |
One test script then works for every row, because the expected result travels with the data. Adding a negative case means adding a row, not a request.
The example API is a hypothetical POST /users signup endpoint with
three rules: email must be a valid address, plan must be free,
pro, or team, and seats must fit the plan (exactly 1 on free, 1-5
on pro, 2-50 on team). The generated file needs:
- Four negative rows, each breaking exactly one rule: an empty email, a malformed email, an unknown plan, and a free plan with extra seats. One defect per row means a 422 can only have one cause.
- Valid rows whose seat count always matches the plan, so the correlation that dynamic variables can't express is built in.
Generate the Postman collection runner data file (JSON) from a template
Here is the template. It writes a JSON array with one object per iteration:
[
<for(i, 1, getParam('rows', 12))><if(getVar('i') > 1)>,<endIf>
<if(getVar('i') == 3)><setVar('plan', 'platinum')>
<elseIf(getVar('i') == 4)><setVar('plan', 'free')>
<else><setVar('plan', getRandomElement('free', 'pro', 'pro', 'team'))><endIf>
{
"case": <if(getVar('i') == 1)>"empty email"
<elseIf(getVar('i') == 2)>"malformed email"
<elseIf(getVar('i') == 3)>"unknown plan"
<elseIf(getVar('i') == 4)>"free plan with extra seats"
<else>"valid"<endIf>,
"name": "<getRandomFullName()>",
"email": <if(getVar('i') == 1)>""
<elseIf(getVar('i') == 2)>"not-an-email"
<else>"<getRandomEmail('example.com')>"<endIf>,
"plan": "<getVar('plan')>",
"seats": <if(getVar('i') == 4)><getRandomNumber(2, 10)>
<elseIf(getVar('plan') == 'pro')><getRandomNumber(1, 5)>
<elseIf(getVar('plan') == 'team')><getRandomNumber(2, 50)>
<else>1<endIf>,
"expectedStatus": <if(getVar('i') <= 4)>422<else>201<endIf>,
"expectedField": <if(getVar('i') <= 2)>"email"
<elseIf(getVar('i') == 3)>"plan"
<elseIf(getVar('i') == 4)>"seats"
<else>""<endIf>
}<end_for>
]
How it works:
- Row count is a parameter.
getParamreadsrowsfrom the request and falls back to 12. The same template can produce a short smoke-test file and a longer regression file. The parameterized templates post goes further with this pattern. planis the only variable, becauseseatsdepends on it. Each row first fixes its plan:platinumon row 3,freeon row 4, and a random pick everywhere else. Listing'pro'twice ingetRandomElementmakes it twice as likely as the other plans, because the function picks one listed value uniformly.planis printed once and read again byseats, which is the only reason it gets asetVar.- Everything else is decided inline, in its own field.
nameis a plain function call.case,email,expectedStatus, andexpectedFieldeach branch on the row index right where the value is written.seatsbranches on the index first (row 4 gets 2-10 seats on a free plan) and then on the plan. Rows 1-4 break one rule each and expect 422; every other row is valid and expects 201. The control flow docs coverforandif/elseIf/else. - Quotes sit inside the branches. Writing
"empty email"rather than wrapping the wholeifin quotes lets each branch go on its own line. The line breaks fall outside every string value, so they only add whitespace between JSON tokens and the output still parses.
The banding is deliberate. Template expressions support comparisons
(==, !=, <, <=, >, >=) and &&/||, but no arithmetic, so
a running "invalid rows so far" counter isn't possible. Comparing the
loop index to fixed values is. The loop end is inclusive, the loop
variable is read with getVar('i'), and the comma guard
<if(getVar('i') > 1)>,<endIf> puts a comma before every row except
the first.
The negative rows come first on purpose. They always occupy iterations
1-4, as long as rows is at least 4, so "iteration 3 failed" always
means the unknown-plan case. Valid rows fill iterations 5 onward.
Rendered with seed 20261002, the first rows look like this (12 rows
in total):
[
{
"case": "empty email",
"name": "Isaiah Craig",
"email": "",
"plan": "free",
"seats": 1,
"expectedStatus": 422,
"expectedField": "email"
},
{
"case": "malformed email",
"name": "Flossie Voltaire",
"email": "not-an-email",
"plan": "pro",
"seats": 5,
"expectedStatus": 422,
"expectedField": "email"
},
{
"case": "unknown plan",
"name": "Belin Holman",
"email": "[email protected]",
"plan": "platinum",
"seats": 1,
"expectedStatus": 422,
"expectedField": "plan"
},
{
"case": "free plan with extra seats",
"name": "Alina Pham",
"email": "[email protected]",
"plan": "free",
"seats": 7,
"expectedStatus": 422,
"expectedField": "seats"
},
{
"case": "valid",
"name": "Morgan Perkins",
"email": "[email protected]",
"plan": "team",
"seats": 39,
"expectedStatus": 201,
"expectedField": ""
}
]
A few dozen rows is plenty for a collection run, since every row is a real HTTP request. The template is nowhere near the engine's per-generation limits (2 seconds, 100,000 node evaluations, 10,000 loop iterations, 8 MiB of output). The limit you'll notice first is how long the run takes.
Save the API response as your Postman data file
Save the template once with
POST /v1/templates. Keeping the
template body in a file and wrapping it with jq avoids escaping it by
hand. The response is the saved template, including its templateId:
jq -n --rawfile body postman/signup-cases.tmpl \
'{name: "Signup API cases", body: $body}' \
| curl -sS --fail -X POST https://api.jsonfabrica.com/v1/templates \
-H "Authorization: Bearer $JSONFABRICA_API_KEY" \
-H "Content-Type: application/json" \
-d @- \
| jq -r '.templateId'
Then generate the data file with
POST /v1/templates/{templateId}/generate. The response has data,
the generated document already parsed as JSON, and meta, whose seed
echoes the seed that was used. The Postman data file is just data:
#!/usr/bin/env bash
# scripts/postman-data.sh: regenerate the signup data file on demand.
set -euo pipefail
API=https://api.jsonfabrica.com/v1
SEED=20261002 # change only when you mean to replace the data
ROWS=12
# The templateId returned by POST /v1/templates above.
TEMPLATE_ID=${TEMPLATE_ID:?set to the templateId from POST /v1/templates}
OUT=postman/data/signup-cases.json
curl -sS --fail -X POST "$API/templates/$TEMPLATE_ID/generate" \
-H "Authorization: Bearer $JSONFABRICA_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"seed\": $SEED, \"params\": {\"rows\": $ROWS}}" \
-o response.json
# Stop if the seed wasn't applied, then keep only the iteration array.
jq -e --argjson s "$SEED" '.meta.seed == $s' response.json > /dev/null
jq '.data' response.json > "$OUT"
rm response.json
echo "$OUT: $(jq length "$OUT") iterations"
The seed check matters. If a typo drops the seed, the API picks a random one, and without the check you'd commit a file nobody can regenerate.
If you'd rather not store the template in JsonFabrica at all,
POST /v1/templates/generate takes the raw template as body along
with the same seed and params, and saves nothing:
jq -n --rawfile body postman/signup-cases.tmpl \
'{body: $body, seed: 20261002, params: {rows: 12}}' \
| curl -sS --fail -X POST https://api.jsonfabrica.com/v1/templates/generate \
-H "Authorization: Bearer $JSONFABRICA_API_KEY" \
-H "Content-Type: application/json" \
-d @- \
| jq '.data' > postman/data/signup-cases.json
Commit the data file next to the exported collection. Every run, local or CI, then reads the same file from disk. CI needs no JsonFabrica API key and makes no call to JsonFabrica, and the data only changes when someone reruns the script and commits the result. For more on wiring generation into scripts and pipelines in general, see API-first data generation.
Use pm.iterationData to assert the expected status
The request body references the data file's fields as variables:
{
"name": "{{name}}",
"email": "{{email}}",
"plan": "{{plan}}",
"seats": {{seats}}
}
{{seats}} is unquoted so it substitutes as a JSON number. Postman's
body editor may flag it as invalid JSON. The request it sends is still
valid once the value is substituted.
The test script reads the expected outcome from the same row. In recent Postman versions it goes in the request's Post-response script; in older versions, the Tests tab:
const expected = pm.iterationData.get('expectedStatus');
const label = pm.iterationData.get('case');
const n = pm.info.iteration + 1; // pm.info.iteration starts at 0
pm.test(`#${n} ${label}: returns ${expected}`, () => {
pm.response.to.have.status(expected);
});
if (expected === 422) {
const field = pm.iterationData.get('expectedField');
pm.test(`#${n} ${label}: flags ${field}`, () => {
// Adjust to your API's error format.
const fields = (pm.response.json().errors || []).map((e) => e.field);
pm.expect(fields).to.include(field);
});
}
if (expected === 201) {
pm.test(`#${n} ${label}: stores plan and seats`, () => {
const user = pm.response.json();
pm.expect(user.plan).to.eql(pm.iterationData.get('plan'));
pm.expect(user.seats).to.eql(pm.iterationData.get('seats'));
});
}
Three things make this script hold up:
- No per-case logic. The script never decides whether a row should
pass. The row says so, and the script compares it with
pm.response.code. A new negative case needs a template change, not a script change. - A 422 must name the right field. A status-only check passes when
the API rejects the request for the wrong reason. Checking
expectedFieldcatches a 422 for "seats" on the row that was meant to test the email. - Test names carry the case. The runner and Newman show
#3 unknown plan: returns 422, so a failure report names the case without anyone opening the data file.
Sending the request on its own with Send, outside a run, leaves
pm.iterationData empty, so these tests fail with undefined. That's
expected. The request is built to run from the data file.
Some negative cases can't be expressed as flat strings, such as a
missing key or a null. For those, have the template write a nested
body object per row and leave the key out of that row's object. Then
send it from a pre-request script with
pm.variables.set('body', JSON.stringify(pm.iterationData.get('body')))
and set the raw request body to {{body}}. {{...}} substitution
doesn't serialize objects, so the JSON.stringify step is required.
Run the data file in the Postman Collection Runner
Locally, open the collection in the Collection Runner and select
postman/data/signup-cases.json as the data file in the run
configuration. Postman sets the iteration count from the number of rows,
and its preview shows the rows it parsed, which is a quick check that
the file loaded as JSON and not as one big string.
The run produces one result per iteration: four 422s at the top, then the 201s. Because the file is committed, a colleague who pulls the branch and runs it sees the same names, emails, and seat counts, and gets the same failures.
Newman data file in CI: newman run -d
Newman is Postman's command-line collection runner, distributed as an
npm package. Export the collection as Collection v2.1 JSON, commit it
next to the data file, and pass the data file with -d, or its long
form --iteration-data:
npx newman run postman/signup.postman_collection.json \
-d postman/data/signup-cases.json \
--env-var "baseUrl=http://localhost:3000" \
-r cli,junit \
--reporter-junit-export results/newman-signup.xml
What each part does:
-druns one iteration per row. Don't pass-nas well; the row count already sets the number of iterations.--env-varsetsbaseUrlfor this run, so the collection's{{baseUrl}}/userspoints at the service your CI job started. Use-e environment.jsoninstead if you already export environments.-r cli,junitprints the usual console output and writes a JUnit XML report that most CI systems can display.
Newman exits with a non-zero status when any test fails, so the CI step
fails without extra scripting. Add --bail if you'd rather stop at the
first failure than see every failing case in one run.
This is where the fixed seed pays off. When #11 valid: returns 201
fails in CI, iteration 11 in the committed file holds the exact values
that failed. Run the Collection Runner locally with that file, or run
Newman against your local server, and you get the same request with the
same body.
When Postman dynamic variables are the right tool
Dynamic variables are still the right tool when a value only has to be
different each run and nobody will assert on it: a random title for a
draft you'll delete, a {{$guid}} idempotency key, a
{{$timestamp}} in a note field. If replaying the exact value wouldn't
help you debug anything, a data file is unnecessary ceremony.
They're also the fix for a real problem with seeded data: unique constraints. If your API rejects a duplicate email with a 409 and the test database survives between runs, the second run against the same database fails on every valid row. Resetting the database per CI run is the cleanest fix. If you can't, keep the seeded email and append a per-run tag in a pre-request script:
// Pre-request: make valid emails unique per run; leave invalid ones.
if (pm.iterationData.get('expectedStatus') === 201) {
const email = pm.iterationData.get('email');
const tag = pm.variables.replaceIn('{{$timestamp}}');
pm.variables.set('email', email.replace('@', `+${tag}@`));
}
pm.variables.set creates a local variable, which takes precedence
over the data file's value, so {{email}} in the body picks up the
tagged address. Only the uniqueness part varies between runs. Every
value the tests actually check still comes from the seeded file.
So the split is: the data file holds anything that defines a test case or appears in an assertion, and dynamic variables handle uniqueness and filler. A collection can use both.
Changing the data without losing reproducibility
A seed fixes the sequence of random draws, so the same template, seed, and parameters return the same file every time. A few changes are worth understanding before you make them:
- Raising
rowsadds rows at the end. The row count doesn't consume any randomness, so with seed20261002,rows: 40returns the same first 12 rows asrows: 12, plus 28 new ones. - Editing the template usually reshuffles the data. Add or remove a random call and every value drawn after that point shifts. That's fine, as long as the regenerated file is committed in the same pull request as the template change and reviewed with it.
- Changing
SEEDdeliberately gives you a fresh data set. It's just as reproducible as the old one.
To find inputs you didn't plan for, run a separate nightly job that
omits the seed. JsonFabrica then picks one and returns it in
meta.seed. Log that value next to the Newman results. When the
nightly run finds a failure, regenerate with that seed to get the exact
file back, and add the case to the committed data.
Reproducible test data from seeds
covers this workflow in more depth.
FAQ
How do I use a data file in Postman?
Open the collection in the Collection Runner, choose your JSON or CSV
file in the data file option of the run configuration, and start the
run. Postman runs the collection once per row, and each row's keys
become variables you can reference as {{name}} in requests or read
with pm.iterationData.get('name') in scripts. From the command line,
Newman does the same with newman run collection.json -d data.json.
What format does a Postman data file need?
Either a CSV file with a header row, or a JSON file containing an array
of objects, one object per iteration. JSON is the better choice for test
cases because the types are in the file: 201 is a number and true is
a boolean. With CSV, types are guessed: Newman converts unquoted numbers
but leaves true as the string "true". Each object's keys become the
variable names for that iteration.
How do I pass a data file to Newman?
Use the -d flag, whose long form is --iteration-data:
newman run collection.json -d data.json. Newman runs one iteration per
row in the file and exits with a non-zero code if any test fails, so a
CI job fails on its own. Add -r cli,junit with
--reporter-junit-export to write a JUnit report for your CI system.
How do I access data file variables in Postman scripts?
In request fields such as the URL, headers, and body, reference them as
{{variableName}}. In pre-request and post-response scripts, read them
with pm.iterationData.get('variableName'), which returns the value for
the current iteration. pm.info.iteration gives the zero-based
iteration index, which is useful in test names.
Can you seed Postman dynamic variables like $randomFirstName?
Postman doesn't document a way to seed its dynamic variables. Each
{{$randomFirstName}} or {{$randomEmail}} reference resolves to a new
random value every time it's used, so a failing run can't be replayed
with the same values. For reproducible runs, generate the values into a
data file from a fixed seed and commit that file.
Does JsonFabrica have a Postman or Newman integration?
No. JsonFabrica is an HTTP API that returns generated JSON. It has no
Postman collection, no Postman app integration, and no Newman plugin.
You call the API with curl or a script, save the data field of the
response as your data file, and point the Collection Runner or
newman run -d at that file.
One template, one fixed seed, and one committed file give your collection valid and negative cases that run the same way in the Collection Runner and in Newman. Seeded, parameterized generation is part of the JsonFabrica API, and the templates API reference documents the request and response.
Generate realistic test data with JsonFabrica
Describe the shape of your data once, then generate as many fresh, realistic JSON documents as you need via a simple API call.