scraping-tweets-from-account

v2026.09.24

Scrapes all tweets, replies, and media from any Twitter/X account using apidojo's Tweet scraper on Apify. Triggers when the user asks to: get all tweets from a Twitter account, export a user's tweet history, scrape a specific Twitter profile's posts, fetch the latest tweets from an account, download tweet data from a user timeline, or collect all posts from a Twitter username. Returns tweet text, likes, retweets, replies, media URLs, and timestamp per tweet. Ideal for journalists, researchers, competitive analysts, and data engineers.

GitHub
安装命令
npx skhub add apidojo-io/scraping-tweets-from-account
Markdown
SKILL.md

Scraping Tweets from an Account

Exports the full tweet history of any public Twitter/X account. Returns the raw timeline dataset including replies and media.

Prerequisites

  • APIFY_TOKEN environment variable set
  • Optional: Apify MCP server installed

Inputs

ParameterTypeRequiredDefaultNotes
searchTermsarray✅[]Twitter advanced search queries (e.g. ["#AI lang:en", "from:NASA"])
sortstringOptionalTopSort order: Latest, Top, or Latest+Top
tweetLanguagestringOptional—ISO 639-1 language code (e.g. en)
maxItemsnumberOptionalUnlimitedMaximum tweets to return
onlyVerifiedUsersbooleanOptionalfalseOnly tweets from verified users
onlyTwitterBluebooleanOptionalfalseOnly Twitter Blue subscribers
onlyImagebooleanOptionalfalseOnly tweets with images
onlyVideobooleanOptionalfalseOnly tweets with videos
onlyQuotebooleanOptionalfalseOnly quote tweets
authorstringOptional—Filter to a specific author handle
inReplyTostringOptional—Tweets replying to a specific handle
mentioningstringOptional—Tweets mentioning a specific handle
geotaggedNearstringOptional—Tweets near a location
withinRadiusstringOptional—Radius around geotaggedNear
geocodestringOptional—Lat/lng + radius string
placeObjectIdstringOptional—Tweets tagged with a place
minimumRetweetsnumberOptional—Minimum retweet count
minimumFavoritesnumberOptional—Minimum like count
minimumRepliesnumberOptional—Minimum reply count
startstringOptional—Tweets after this date (YYYY-MM-DD)
endstringOptional—Tweets before this date (YYYY-MM-DD)
includeSearchTermsbooleanOptionalfalseAdd the matched search term to each tweet
customMapFunctionstringOptional—JavaScript function to transform each output object

Workflow

Progress:
- [ ] Step 1: Confirm account is public
- [ ] Step 2: Run tweet-scraper
- [ ] Step 3: Poll for SUCCEEDED
- [ ] Step 4: Fetch and deliver dataset

Step 1: Validate Input

Strip @ from username if present. Do not attempt to scrape private accounts — the actor will return 0 results.

Step 2: Run the Actor

Recommended — run_actor.js (handles waiting, output, and file saving automatically):

# Quick answer (prints table to chat)
node scripts/run_actor.js \
  --actor "apidojo~tweet-scraper" \
  --input '{"param": "value"}'

# Save as CSV
node scripts/run_actor.js \
  --actor "apidojo~tweet-scraper" \
  --input '{"param": "value"}' \
  --output YYYY-MM-DD_results.csv --format csv

# Save as JSON
node scripts/run_actor.js \
  --actor "apidojo~tweet-scraper" \
  --input '{"param": "value"}' \
  --output YYYY-MM-DD_results.json --format json

APIFY_TOKEN must be set in environment or .env file.

If Apify MCP is available:

Tool: apify:run-actor
Actor: "apidojo~tweet-scraper"
Input:
{
  "twitterHandles": ["<handle>"],
  "maxItems": 200,
  "sort": "Latest"
}

REST API fallback:

curl -X POST \
  "https://api.apify.com/v2/acts/apidojo~tweet-scraper/runs?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"twitterHandles": ["<handle>"], "maxItems": 200, "sort": "Latest"}'

Save id as RUN_ID. Poll until status = SUCCEEDED:

curl "https://api.apify.com/v2/actor-runs/$RUN_ID?token=$APIFY_TOKEN" | grep '"status"'

Fetch results:

curl "https://api.apify.com/v2/actor-runs/$RUN_ID/dataset/items?token=$APIFY_TOKEN&format=json"

Step 3: Handle Edge Cases

  • 0 results: Account may be private, suspended, or handle misspelled. Report to user.
  • Fewer results than expected: Account may have fewer public tweets than requested. Return what is available.
  • Protected account: Actor returns empty — inform user the account is private.

Output Format

# Tweet Timeline: @<username>
Tweets collected: N | Includes replies: YES/NO | Includes retweets: YES/NO

| Tweet ID | Text (truncated) | Likes | Retweets | Replies | Media | Timestamp |
|----------|-----------------|-------|----------|---------|-------|-----------|
| ...      | ...             | ...   | ...      | ...     | ...   | ...       |

Full dataset: N rows × 12 fields
Available fields: id, text, likeCount, retweetCount, replyCount, quoteCount,
isReply, isRetweet, media, tweetUrl, lang, createdAt

Troubleshooting

0 results for a known public account: Try again — Twitter may rate-limit intermittently. Missing older tweets: Twitter API limits historical access; very old tweets may not be available. Timeout: Reduce maxItems to 500 max per run.

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

Apache-2.0

源路径

skills/primitives/scraping-tweets-from-account

默认分支

main

最新提交

ffbdc00

Tree SHA

c7df562