Export Data via API: A Full Dump the Export Button Misses
Suppose you export data via the API because the export button left things out. The script points
at HubSpot's contact search endpoint, asks for 200 records at a time, and follows the after value
from page to page. If the account holds, say, 14,212 contacts, the full dump runs cleanly for fifty
pages and then the fifty-first request comes back as a 400. Nothing is wrong with the script. HubSpot's own search documentation says, in the paragraph most people skip on the way to the
code sample, that "the search endpoints are limited to 10,000 total results for any given query.
Attempting to page beyond 10,000 will result in a 400 error"
(HubSpot CRM Search API, read 19 September 2026).
So the dump has exactly 10,000 contacts in it, which is a round number, which is the only reason
anyone notices.
That is the general shape of a one-time API dump going wrong. The export button in the app failed you for a reason covered elsewhere on this site (flattened relations, a view filter, links instead of files; see why a CSV export is not a backup), so you went to the API, which feels like the raw source. It is closer. But an API is built for integrations that read a few hundred records at a time, and a full dump runs into every ceiling that everyday use never touches: result caps on search, default field sets, records the list endpoint excludes on purpose, rate limits, and tokens that see less than an admin sees.
The examples below come from three vendors whose developer documentation I read on 19 September 2026: HubSpot (CRM records), Asana (tasks and comments) and Zendesk (support tickets). Your vendor will differ in the numbers. The categories of failure repeat almost everywhere.
Walk the list endpoint, not the search endpoint
Search endpoints are the ones a script reaches for first, because they accept filters and sort orders. They are also the ones with a hard total.
| Vendor and endpoint | Per page | Total a single query can reach | What happens past it |
|---|---|---|---|
| HubSpot CRM search | up to 200 | 10,000 | 400 error |
HubSpot list, GET /crm/objects/.../contacts |
up to 100 | no cap stated; follow after |
— |
| Zendesk Search API | up to 100 | 1,000 | 422 at page 11 |
| Zendesk offset pagination on list endpoints | 100 | 100 pages, 10,000 resources | 400 |
| Zendesk incremental export (cursor) | up to 1,000 | none; runs until end_of_stream |
— |
| Asana paginated collections | 1 to 100 | page through all | unpaginated calls truncate near 1,000 |
Zendesk's search reference is blunt about it: the API "returns up to 1,000 results per query, with a maximum of 100 results per page," and requesting page 11 at 100 per page returns "a 422 Insufficient Resource Error." The same page points people with large datasets at the incremental export endpoints instead (Zendesk Search API, read 19 September 2026). The general pagination page adds the second trap. Offset pagination, which is what you get unless you opt in to cursors, "is limited to the first 100 pages and 10,000 resources," and anything past that returns a 400 whose body explains how to switch to cursor pagination (Zendesk pagination, read 19 September 2026).
Asana's version is quieter. Its pagination guide says an unpaginated query "will be truncated at
about 1,000 objects," and that on large organisations it may simply time out, so the limit
parameter (1 to 100) is not optional for a dump
(Asana pagination, updated 12 September 2025, read 19 September 2026).
The search endpoints are not useless. They are, in fact, the best tool for the last step of this
job, which is counting. Zendesk notes that the count property "shows the actual number of results.
For example, if a query has 5,000 results, the count value will be 5,000, even if the API only
returns the first 1,000 results." HubSpot's search response carries a total in the same way. Use
the search endpoint to learn how many records there should be, and the list or incremental endpoint
to fetch them.
The default response is a summary of the record
The second ceiling is invisible because nothing errors. You get every record, and every record is missing most of its fields.
Asana returns compact records unless told otherwise. The endpoint that lists tasks in a project
"Returns the compact task records for all tasks within the given project," and the input/output
options guide explains that opt_fields is how you "list the exact set of fields that the API
should return for the objects"
(Asana input/output options, read 19 September 2026).
The compact task is little more than an ID and a name. Assignee, due date, notes, custom field
values, memberships, parent, dependencies: each has to be named in opt_fields, and the opt_fields
list in the reference example for that one endpoint has 115 entries. Comments are
not in it at all. They are stories, fetched from a separate endpoint per task, and that endpoint
also returns compact records by default.
HubSpot returns the properties you name. Its search documentation lists the default for
contacts: createdate, email, firstname, hs_object_id, lastmodifieddate and lastname.
Six fields, in an account that may have two hundred. The list endpoint takes a comma-separated
properties parameter, and its reference says that requested properties "not present on the
requested object(s)... will be ignored," so a field missing from the dump tells you nothing about
whether it was empty or never asked for. Its limit also defaults to 10 per page, so set it
explicitly
(HubSpot list contacts reference, read 19 September 2026).
To build the properties parameter,
fetch the property definitions first with GET /crm/properties/2026-03/{objectType}, which returns
every property defined for the object, custom ones included
(HubSpot properties API, read 19 September 2026).
Store that response too. It is the schema, with field types and dropdown options, and it is the
only thing that will tell someone in two years what deal_stage_custom_7 meant. HubSpot also
offers propertiesWithHistory, which returns previous values alongside current ones and, per the
same reference, reduces "the maximum number of contacts that can be read by a single request."
Decide whether the change history matters to you before the dump, because it is not recoverable
after, and budget the extra pages if it does.
Zendesk returns tickets without their conversation. The incremental ticket export gives you
ticket objects. The comments live in a different stream, the incremental ticket event export, and
even there the documentation says that without the comment_events sideload "any comment present
in the ticket update is described only by Boolean comment_present and comment_public object
properties... The comment itself is not included"
(Zendesk incremental exports, read 19 September 2026).
A support archive with every ticket and none of the replies is a list of subject lines.
The practical rule that falls out of all three: before writing the loop, fetch one record through the list endpoint, one through the single-record endpoint with every field requested, and diff them. Whatever is in the second and not in the first is what your dump would have lost.
What the list endpoint leaves out on purpose
Some records are excluded by design, and the exclusion is documented, just not where you look.
- Archived HubSpot records. The list endpoint's
archivedparameter is described as "Whether to return only results that have been archived." Only. So a single pass with the default gives you active records, and archived ones need a second full pass witharchived=true. Archived records also "won't appear in any search results," which means your search-based count will not include them either. - Archived Zendesk tickets. The tickets reference states that archived tickets are not included
in the responses from
GET /api/v2/tickets,GET /api/v2/organizations/{organization_id}/ticketsorGET /api/v2/tickets/recent, and then says: "To get a list of all tickets in your account, use the Incremental Ticket Export" (Zendesk tickets API, read 19 September 2026). Zendesk's help centre says that "in most cases" it archives tickets automatically "120 days after the ticket status changes to Closed," and sooner for accounts with extremely high ticket volumes (About ticket archiving, read 19 September 2026), so in an older account the ordinary list endpoint can return a small fraction of the real history. - Deleted Zendesk tickets, going the other way. In the incremental export, "By default,
deletions will appear in the ticket stream." If you are building an archive of what customers
actually saw, set
exclude_deleted=trueor filter afterwards. If you are building an audit trail, keep them. - Asana subtasks. The project task list returns the tasks in the project. Subtasks have their
own endpoint,
GET /tasks/{task_gid}/subtasks, called one parent at a time, and a subtask carries its ownprojectsandmembershipsfields rather than inheriting its parent's. The reference does not promise that the project list includes subtasks, so do not count on it; fetch them through the subtasks endpoint and de-duplicate by ID. Ask fornum_subtasksinopt_fieldson the first pass and you only need to call the subtasks endpoint for the tasks where it is above zero, which usually saves most of those requests. - Attachments, everywhere. Asana's attachments endpoint returns compact attachment records, and
its bulk export states that attachment objects "do not include
download_urlorview_url," so the files have to be fetched from the attachments endpoint for current URLs. The clock on those URLs is the problem covered in downloading attachments from your export: fetch the bytes during the dump, not from the links afterwards.
Rate limits decide how long this takes, so do the arithmetic first
Every one of these vendors returns 429 Too Many Requests when a limit is hit, with a
Retry-After header that says how many seconds to wait. The numbers underneath differ by an order
of magnitude, and on some of them by plan.
| Vendor | Limit relevant to a dump | Source |
|---|---|---|
| Asana | 150 requests per minute on a free domain, 1,500 on a paid one; search 60 per minute; 50 concurrent GETs | Asana rate limits |
| HubSpot, privately distributed app | Free and Starter: 100 per 10 seconds per app, 250,000 per day per account. Professional: 190 and 625,000. Enterprise: 190 and 1,000,000 | HubSpot usage guidelines |
| HubSpot search | 5 requests per second per account | HubSpot CRM search page |
| Zendesk incremental export | 10 requests per minute | Zendesk incremental exports page |
| Zendesk Search API | 100 requests per minute per account, also counted against the global limit | Zendesk search page |
| Zendesk Suite, general API | 200 (Team), 400 (Growth and Professional), 700 (Enterprise), 2,500 (Enterprise Plus) per minute | Zendesk rate limits |
All figures read on 19 September 2026. Plan names and numbers change; recheck the linked page on the day you run the dump.
Now the arithmetic, which is the part worth doing on paper before any code exists.
Take an Asana project with 4,000 tasks. Listing them at 100 per page is 40 requests. Comments are one
stories request per task: 4,000. Attachments, one request per task: another 4,000. Subtasks, if you
skip the num_subtasks shortcut: 4,000 more. Call it 12,040 requests. On a free domain at 150 a
minute that is a little over 80 minutes of perfectly paced calls. On a paid domain at 1,500 a
minute, about 8 minutes. Same data, ten times the wall-clock time, decided entirely by the plan
you are paying for, which matters if you are planning to downgrade before cancelling. Run the dump
first.
Zendesk's incremental ticket export looks slow at 10 requests a minute until you notice the page size: up to 1,000 tickets per page, so the ticket list moves at up to 10,000 tickets a minute. A 60,000-ticket account lists in about six minutes. The comment stream is the long part, because events outnumber tickets many times over.
Two rules from Asana's rate-limit page apply to every vendor, whether or not they write them down.
First, "Requests rejected by this limiter still count against your quota," so a script that hammers
through 429s makes itself slower. Second, the quota "is evaluated more frequently than once per
minute," so Retry-After is often under 60 seconds; "always use the value returned, never a
hard-coded minute"
(Asana rate limits, updated 3 July 2026, read 19 September 2026).
HubSpot adds a condition that matters if you share the account with live integrations: the daily
limit on private apps "is shared across all apps within the same HubSpot account." A dump that burns
a Starter account's 250,000 daily calls by mid-afternoon also stops the accounting sync for the rest
of the day.
Whose eyes the token has
An API dump sees exactly what its credential is allowed to see, and the credential is usually a person's.
Asana's personal access token documentation puts it plainly: "Your tokens act on your behalf when interacting with the API." A token generated by a project lead will not return private projects the lead was never added to. The dump will look complete, and it will be complete for that person. Generate it from the account with the widest access, and check the project count against what an admin sees in the app.
Zendesk's incremental export endpoints are marked Allowed For Admins, every one of them. Generate the API token from an admin account, and try one request against the export endpoint before planning around it.
Expiry is the other half. Asana's OAuth guide says access tokens "expire in 1 hour (3600 seconds)";
a dump that runs for 80 minutes on an OAuth token dies at minute 60 unless the script refreshes it.
Personal access tokens are described as "persistent by default," which makes them the simpler choice
for a one-off job, with the obvious condition that you revoke the token when the dump is done.
HubSpot's guidelines say generated access tokens carry an expires_in value and that apps are
responsible for refreshing them. Check the expiry on whichever credential you use against your
arithmetic from the last section, before starting.
Write each page to disk exactly as it arrived
The instinct is to transform while fetching: pull a page, flatten it, append to a spreadsheet. Do not. Save the raw response body, one page per line or one file per page, and transform later from the saved copy. Three reasons.
The raw JSON keeps every field, including the ones you did not know you needed. A transformation written on day one encodes day-one assumptions, and the vendor account may not exist on day forty when the assumption turns out wrong.
The saved pages make the job resumable. Zendesk's cursor export returns an after_url and an
end_of_stream flag; save the cursor after each page is safely written and a crash at page 300
restarts at page 301. Be careful which positions are safe to save. Asana warns that "an offset token
will expire after some time, as data may have changed," so for Asana the durable resume point is
the last task ID you finished, not the offset string.
And the raw pages are what the archive should hold anyway, for the reasons in keeping a retired system readable.
A minimal loop against Zendesk's cursor-based ticket export, as a sketch rather than a tool, looks like this:
import json, os, time, requests
START = ("https://SUBDOMAIN.zendesk.com/api/v2/incremental/tickets/cursor"
"?start_time=0&per_page=1000")
auth = ("admin@example.com/token", os.environ["ZENDESK_API_TOKEN"])
url = START # resume from the saved cursor if a previous run stopped part way
if os.path.exists("cursor.txt"):
with open("cursor.txt", encoding="utf-8") as f:
url = f.read().strip() or START
with open("tickets.ndjson", "a", encoding="utf-8") as out:
while url:
r = requests.get(url, auth=auth, timeout=60)
if r.status_code == 429:
time.sleep(int(r.headers.get("Retry-After", "60")))
continue
r.raise_for_status()
page = r.json()
out.write(json.dumps(page) + "\n")
out.flush()
os.fsync(out.fileno())
with open("cursor.txt", "w", encoding="utf-8") as f:
f.write(page.get("after_url") or "")
url = None if page.get("end_of_stream") else page.get("after_url")
It writes each page before recording the cursor, picks up from that cursor if it is run again,
waits the time the server asks for, and stops on the flag the vendor provides rather than on an
empty page. The token comes from an environment variable so it never sits in the script or its
history. The same habits carry over to any vendor. The start_time=0 asks for everything from the beginning; Zendesk requires a start time for
the first request, says it "must be more than one minute in the past to avoid missing data," and does not return data for the most recent minute.
Check the vendor's own bulk export before writing anything
Some vendors have built a proper dump and put it behind a plan tier, which is worth knowing before spending a day on a script.
Asana has three. Its resource export bulk-exports tasks, teams and messages for a workspace as "JSON Lines format and compressed in a gzip container," with attachments and stories included on tasks by default. The catches are all in the reference: it is available only to service accounts "from an Enterprise+ organization or an organization with the Compliance Management Add-on," a workspace can have "one in progress export request at a given time," the file expires "30 days after its completion," and "Exports currently include undeleted objects." The graph export, starting from a project, team or portfolio, is Enterprise and Enterprise+ only, and when it covers more than 1,000 tasks its result is cached for four hours, so running it twice within that window returns the same file. The organization export is for Enterprise service accounts (Asana exports reference, read 19 September 2026).
If your plan qualifies, use it, and treat the script from this page as the check on it rather than the source. If it does not, the plan gate is itself a piece of information: an upgrade for one month may cost less than the hours, and it goes on the list of things to price in the twelve things to capture before you cancel.
Counting what came back
The dump is finished when three numbers agree, not when the loop ends.
- The vendor's count. A search query with no filter returns a total (HubSpot's
total, Zendesk'scount), even when it will not let you page past the cap. Asana's task-count endpoint for a project returnsnum_tasks,num_incomplete_tasksandnum_completed_tasks, though only the fields you opt into and at a high cost against its rate limits, so call it once per project. - The count in your raw files. Count records in the saved pages, de-duplicated by ID. Zendesk describes the incremental ticket export as returning "the tickets that changed since the start time," and the reference does not promise that each ticket appears only once, so a ticket that changes while the dump runs may come back on a later page. Keep the most recent copy per ID rather than treating a repeat as an error.
- The count after transformation. If the spreadsheet or database you built has fewer rows than the raw files, the loss happened in your code, and the raw files mean you can fix it without going back to the vendor.
Then spot-check depth, not only breadth. Pick five records with long histories, open them in the app, and compare comment counts, attachment counts and custom field values against the dump. A dump where every count matches and every ticket has zero comments is the failure described in the second section, and the counts alone will never show it.
The difference between the search cap and the list endpoint, or between archived=false and the
whole account, is the kind of thing that only surfaces when someone needs the record that is not
there. Doing the count while the account is still live is the whole point of doing it now.
Verified against HubSpot, Asana and Zendesk developer documentation on 19 September 2026, at the pages linked above. Rate limits, page sizes and plan gates are the parts most likely to change; recheck the linked pages on the day you run a dump. This page does not recommend any of these products or compare them; they are used here because their API documentation states its limits precisely enough to quote.
Frequently asked questions
Why does my API export stop at exactly 10,000 or 1,000 records?
Because you are paging through a search endpoint, and search endpoints are capped. HubSpot's CRM search documentation says the search endpoints are limited to 10,000 total results for any given query and that paging beyond 10,000 returns a 400 error. Zendesk's Search API returns up to 1,000 results per query, and asking for page 11 at 100 per page returns a 422. Zendesk's offset pagination on list endpoints has its own ceiling of 100 pages and 10,000 resources. The fix is to switch to the endpoint the vendor built for walking everything: HubSpot's list endpoint with the after cursor, Zendesk's incremental export, or cursor pagination wherever it is offered.
How long does a full API export take?
Work it out from three numbers before you start: records per page, requests allowed per minute, and how many extra requests each record needs. Zendesk's incremental ticket export returns up to 1,000 tickets per page at 10 requests per minute, so the ticket list itself moves at up to 10,000 tickets a minute. Asana returns at most 100 objects per page, but comments live on a separate stories endpoint that takes one request per task, so a 4,000-task project with stories, attachments and subtasks fetched per task is roughly 12,000 requests: about 80 minutes at the 150-per-minute free-domain limit, about 8 minutes at the 1,500-per-minute paid limit.
Does an API export include deleted and archived records?
Only if you ask for them, and each vendor asks differently. HubSpot's list endpoints take an archived parameter that returns only archived records, so archived records need a second pass. Zendesk's GET /api/v2/tickets does not include archived tickets at all, and its documentation sends you to the incremental ticket export for a list of all tickets; that export includes deletions by default unless you set excludedeleted. Asana's bulk resource export states that it currently includes undeleted objects only. Decide which of these you need before the dump, not after the account closes.
What should I do when the API returns 429 Too Many Requests?
Wait for the number of seconds in the Retry-After header, then retry the same request. Asana, Zendesk and HubSpot all document a 429 response for rate limits, and Asana adds that rejected requests still count against your quota, so retrying early pushes recovery further out. Do not hard-code a sleep of one minute either; Asana notes its quota is evaluated more often than once a minute, so the header is usually shorter and always more accurate.