Datasets:
question_id int64 1 3.24k | task_id stringlengths 3 77 | difficulty stringclasses 3
values | problem_description stringlengths 193 3.69k | canonical_tests dict | public_tests listlengths 1 5 | hints listlengths 0 8 | metadata dict | language stringclasses 9
values | interface dict |
|---|---|---|---|---|---|---|---|---|---|
1 | two-sum | Easy | You are given an array of integers nums and an integer target, return indices of the two numbers such that they add up to target.
You may assume that each input would have exactly one solution, and you may not use the same element twice.
You can return the answer in any order.
Example 1:
Input: nums = [2,7,11,15], ... | {
"language": "python",
"source": "def answers_match(result, expected):\n \"\"\"The answer may list its items in any order (only the top-level order is free).\"\"\"\n if isinstance(result, list) and isinstance(expected, list):\n return sorted(result, key=repr) == sorted(expected, key=repr)\n return ... | [
"[2,7,11,15]\n9",
"[3,2,4]\n6",
"[3,3]\n6"
] | [
"A really brute force way would be to search for all possible pairs of numbers but that would be too slow. Again, it's best to try out brute force solutions just for completeness. It is from these brute force solutions that you can come up with optimizations.",
"So, if we fix one of the numbers, say x, we have to... | {
"canonical_parameter_names": [
"nums",
"target"
],
"leetcode_dataset_entry_point": "Solution().twoSum",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:26",
"tests_total": 80,
"t... | cpp | {
"raw_signature": "vector<int> twoSum(vector<int>& nums, int target) {",
"container": "Solution",
"callable": "twoSum",
"parameters": [
{
"name": "nums",
"type": "vector<int>&"
},
{
"name": "target",
"type": "int"
}
],
"return_type": "vector<int>",
"signature_sourc... |
3 | longest-substring-without-repeating-characters | Medium | Given a string s, find the length of the longest substring without duplicate characters.
Example 1:
Input: s = "abcabcbb"
Output: 3
Explanation: The answer is "abc", with the length of 3. Note that "bca" and "cab" are also correct answers.
Example 2:
Input: s = "bbbbb"
Output: 1
Explanation: The answer is "b", with... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(s = \"abcabcbb\") == 3\n assert candidate(s = \"bbbbb\") == 1\n assert candidate(s = \"pwwkew\") == 3\n assert candidate(s = \"abcdabcabcabcd\") == 4\n assert candidate(s = \"abcdefgabcdefgabcdefgabcdefg\") == 7\n assert c... | [
"\"abcabcbb\"",
"\"bbbbb\"",
"\"pwwkew\""
] | [
"There are less than 100 unique characters. We can check all substrings with length at most 100 for example. This is a good enough approximation."
] | {
"canonical_parameter_names": [
"s"
],
"leetcode_dataset_entry_point": "Solution().lengthOfLongestSubstring",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:36",
"tests_total": 69,
"... | cpp | {
"raw_signature": "int lengthOfLongestSubstring(string s) {",
"container": "Solution",
"callable": "lengthOfLongestSubstring",
"parameters": [
{
"name": "s",
"type": "string"
}
],
"return_type": "int",
"signature_source": "leetcode/codeSnippets",
"type_source": "leetcode/codeSnippet... |
4 | median-of-two-sorted-arrays | Hard | Given two sorted arrays nums1 and nums2 of size m and n respectively, return the median of the two sorted arrays.
The overall run time complexity should be O(log (m+n)).
Example 1:
Input: nums1 = [1,3], nums2 = [2]
Output: 2.00000
Explanation: merged array = [1,2,3] and median is 2.
Example 2:
Input: nums1 = [1,2]... | {
"language": "python",
"source": "def answers_match(result, expected):\n \"\"\"Decimal answers are accepted within 1e-5 (absolute or relative), as on LeetCode.\"\"\"\n import math\n if isinstance(result, list) and isinstance(expected, list):\n return len(result) == len(expected) and all(answers_mat... | [
"[1,3]\n[2]",
"[1,2]\n[3,4]"
] | [] | {
"canonical_parameter_names": [
"nums1",
"nums2"
],
"leetcode_dataset_entry_point": "Solution().findMedianSortedArrays",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:38",
"tests_... | cpp | {
"raw_signature": "double findMedianSortedArrays(vector<int>& nums1, vector<int>& nums2) {",
"container": "Solution",
"callable": "findMedianSortedArrays",
"parameters": [
{
"name": "nums1",
"type": "vector<int>&"
},
{
"name": "nums2",
"type": "vector<int>&"
}
],
"re... |
5 | longest-palindromic-substring | Medium | Given a string s, return the longest palindromic substring in s.
Example 1:
Input: s = "babad"
Output: "bab"
Explanation: "aba" is also a valid answer.
Example 2:
Input: s = "cbbd"
Output: "bb"
Constraints:
- 1 <= s.length <= 1000
- s consist of only digits and English letters.
| {
"language": "python",
"source": "def check(candidate):\n assert candidate(s = \"abba\") == \"abba\"\n assert candidate(s = \"aaaa\") == \"aaaa\"\n assert candidate(s = \"abacdfgdcaba\") == \"aba\"\n assert candidate(s = \"ac\") == \"a\"\n assert candidate(s = \"babad\") == \"aba\"\n assert candi... | [
"\"babad\"",
"\"cbbd\""
] | [
"How can we reuse a previously computed palindrome to compute a larger palindrome?",
"If “aba” is a palindrome, is “xabax” a palindrome? Similarly is “xabay” a palindrome?",
"Complexity based hint: If we use brute-force and check whether for every start and end position a substring is a palindrome we have O(n^2... | {
"canonical_parameter_names": [
"s"
],
"leetcode_dataset_entry_point": "Solution().longestPalindrome",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:41",
"tests_total": 124,
"tests_... | cpp | {
"raw_signature": "string longestPalindrome(string s) {",
"container": "Solution",
"callable": "longestPalindrome",
"parameters": [
{
"name": "s",
"type": "string"
}
],
"return_type": "string",
"signature_source": "leetcode/codeSnippets",
"type_source": "leetcode/codeSnippets"
} |
6 | zigzag-conversion | Medium | The string "PAYPALISHIRING" is written in a zigzag pattern on a given number of rows like this: (you may want to display this pattern in a fixed font for better legibility)
P A H N
A P L S I I G
Y I R
And then read line by line: "PAHNAPLSIIGYIR"
Write the code that will take a string and make this conversi... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(s = \"PAYPALISHIRING\",numRows = 4) == \"PINALSIGYAHRPI\"\n assert candidate(s = \"ABCDEFGHI\",numRows = 3) == \"AEIBDFHCG\"\n assert candidate(s = \"A,B,C,D,E,F,G,H,I,J,K,L,M,N,O,P,Q,R,S,T,U,V,W,X,Y,Z\",numRows = 5) == \"AEIMQUY,,... | [
"\"PAYPALISHIRING\"\n3",
"\"PAYPALISHIRING\"\n4",
"\"A\"\n1"
] | [] | {
"canonical_parameter_names": [
"s",
"numRows"
],
"leetcode_dataset_entry_point": "Solution().convert",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:43",
"tests_total": 104,
"t... | cpp | {
"raw_signature": "string convert(string s, int numRows) {",
"container": "Solution",
"callable": "convert",
"parameters": [
{
"name": "s",
"type": "string"
},
{
"name": "numRows",
"type": "int"
}
],
"return_type": "string",
"signature_source": "leetcode/codeSnippe... |
7 | reverse-integer | Medium | Given a signed 32-bit integer x, return x with its digits reversed. If reversing x causes the value to go outside the signed 32-bit integer range [-2^31, 2^31 - 1], then return 0.
Assume the environment does not allow you to store 64-bit integers (signed or unsigned).
Example 1:
Input: x = 123
Output: 321
Example 2... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(x = -2147483412) == -2143847412\n assert candidate(x = 2147483647) == 0\n assert candidate(x = 120) == 21\n assert candidate(x = -123) == -321\n assert candidate(x = 1534236469) == 0\n assert candidate(x = 0) == 0\n ass... | [
"123",
"-123",
"120"
] | [] | {
"canonical_parameter_names": [
"x"
],
"leetcode_dataset_entry_point": "Solution().reverse",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:45",
"tests_total": 61,
"tests_dropped": 1... | cpp | {
"raw_signature": "int reverse(int x) {",
"container": "Solution",
"callable": "reverse",
"parameters": [
{
"name": "x",
"type": "int"
}
],
"return_type": "int",
"signature_source": "leetcode/codeSnippets",
"type_source": "leetcode/codeSnippets"
} |
8 | string-to-integer-atoi | Medium | Implement the myAtoi(string s) function, which converts a string to a 32-bit signed integer.
The algorithm for myAtoi(string s) is as follows:
1. Whitespace: Ignore any leading whitespace (" ").
2. Signedness: Determine the sign by checking if the next character is '-' or '+', assuming positivity if neither present.
... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(s = \"2147483647\") == 2147483647\n assert candidate(s = \"42 with words\") == 42\n assert candidate(s = \"20000000000000000000000000000000000000000\") == 2147483647\n assert candidate(s = \"-2147483649\") == -2147483648\n as... | [
"\"42\"",
"\" -042\"",
"\"1337c0d3\"",
"\"0-1\"",
"\"words and 987\""
] | [] | {
"canonical_parameter_names": [
"s"
],
"leetcode_dataset_entry_point": "Solution().myAtoi",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:47",
"tests_total": 185,
"tests_dropped": 1... | cpp | {
"raw_signature": "int myAtoi(string s) {",
"container": "Solution",
"callable": "myAtoi",
"parameters": [
{
"name": "s",
"type": "string"
}
],
"return_type": "int",
"signature_source": "leetcode/codeSnippets",
"type_source": "leetcode/codeSnippets"
} |
9 | palindrome-number | Easy | Given an integer x, return true if x is a palindrome, and false otherwise.
Example 1:
Input: x = 121
Output: true
Explanation: 121 reads as 121 from left to right and from right to left.
Example 2:
Input: x = -121
Output: false
Explanation: From left to right, it reads -121. From right to left, it becomes 121-. The... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(x = 1221) == True\n assert candidate(x = 10) == False\n assert candidate(x = 123421) == False\n assert candidate(x = 1) == True\n assert candidate(x = -121) == False\n assert candidate(x = 123456) == False\n assert cand... | [
"121",
"-121",
"10"
] | [
"Beware of overflow when you reverse the integer."
] | {
"canonical_parameter_names": [
"x"
],
"leetcode_dataset_entry_point": "Solution().isPalindrome",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:49",
"tests_total": 61,
"tests_droppe... | cpp | {
"raw_signature": "bool isPalindrome(int x) {",
"container": "Solution",
"callable": "isPalindrome",
"parameters": [
{
"name": "x",
"type": "int"
}
],
"return_type": "bool",
"signature_source": "leetcode/codeSnippets",
"type_source": "leetcode/codeSnippets"
} |
10 | regular-expression-matching | Hard | Given an input string s and a pattern p, implement regular expression matching with support for '.' and '*' where:
- '.' Matches any single character.
- '*' Matches zero or more of the preceding element.
Return a boolean indicating whether the matching covers the entire input string (not partial).
Example 1:
Input:... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(s = \"aa\",p = \"a*\") == True\n assert candidate(s = \"aab\",p = \"c*a*b\") == True\n assert candidate(s = \"ab\",p = \".*\") == True\n assert candidate(s = \"aa\",p = \"a\") == False\n assert candidate(s = \"mississippi\",p... | [
"\"aa\"\n\"a\"",
"\"aa\"\n\"a*\"",
"\"ab\"\n\".*\""
] | [] | {
"canonical_parameter_names": [
"s",
"p"
],
"leetcode_dataset_entry_point": "Solution().isMatch",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:52",
"tests_total": 151,
"tests_d... | cpp | {
"raw_signature": "bool isMatch(string s, string p) {",
"container": "Solution",
"callable": "isMatch",
"parameters": [
{
"name": "s",
"type": "string"
},
{
"name": "p",
"type": "string"
}
],
"return_type": "bool",
"signature_source": "leetcode/codeSnippets",
"ty... |
11 | container-with-most-water | Medium | You are given an integer array height of length n. There are n vertical lines drawn such that the two endpoints of the i^th line are (i, 0) and (i, height[i]).
Find two lines that together with the x-axis form a container, such that the container contains the most water.
Return the maximum amount of water a container... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(height = [1, 1]) == 1\n assert candidate(height = [4, 3, 2, 1, 4]) == 16\n assert candidate(height = [8, 10, 14, 0, 13, 10, 9, 9, 8, 9]) == 72\n assert candidate(height = [1, 8, 6, 2, 5, 4, 8, 3, 7]) == 49\n assert candidate(... | [
"[1,8,6,2,5,4,8,3,7]",
"[1,1]"
] | [
"If you simulate the problem, it will be O(n^2) which is not efficient.",
"Try to use two-pointers. Set one pointer to the left and one to the right of the array. Always move the pointer that points to the lower line.",
"How can you calculate the amount of water at each step?"
] | {
"canonical_parameter_names": [
"height"
],
"leetcode_dataset_entry_point": "Solution().maxArea",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:54",
"tests_total": 116,
"tests_dropp... | cpp | {
"raw_signature": "int maxArea(vector<int>& height) {",
"container": "Solution",
"callable": "maxArea",
"parameters": [
{
"name": "height",
"type": "vector<int>&"
}
],
"return_type": "int",
"signature_source": "leetcode/codeSnippets",
"type_source": "leetcode/codeSnippets"
} |
12 | integer-to-roman | Medium | Seven different symbols represent Roman numerals with the following values:
Symbol | Value
I | 1
V | 5
X | 10
L | 50
C | 100
D | 500
M | 1000
Roman numerals are formed by appending the conversions of decimal place values from highest to lowest. Converting a decimal place value into a Roman numeral has the following r... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(num = 44) == \"XLIV\"\n assert candidate(num = 9) == \"IX\"\n assert candidate(num = 4) == \"IV\"\n assert candidate(num = 2023) == \"MMXXIII\"\n assert candidate(num = 589) == \"DLXXXIX\"\n assert candidate(num = 444) == ... | [
"3749",
"58",
"1994"
] | [] | {
"canonical_parameter_names": [
"num"
],
"leetcode_dataset_entry_point": "Solution().intToRoman",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:56",
"tests_total": 77,
"tests_droppe... | cpp | {
"raw_signature": "string intToRoman(int num) {",
"container": "Solution",
"callable": "intToRoman",
"parameters": [
{
"name": "num",
"type": "int"
}
],
"return_type": "string",
"signature_source": "leetcode/codeSnippets",
"type_source": "leetcode/codeSnippets"
} |
13 | roman-to-integer | Easy | Roman numerals are represented by seven different symbols: I, V, X, L, C, D and M.
Symbol Value
I 1
V 5
X 10
L 50
C 100
D 500
M 1000
For example, 2 is written as II in Roman numeral, just two ones added together. 12 is written a... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(s = \"XCIX\") == 99\n assert candidate(s = \"MMCMXCIX\") == 2999\n assert candidate(s = \"MMMCMXCIX\") == 3999\n assert candidate(s = \"DCXXI\") == 621\n assert candidate(s = \"XC\") == 90\n assert candidate(s = \"VIII\") ... | [
"\"III\"",
"\"LVIII\"",
"\"MCMXCIV\""
] | [
"Problem is simpler to solve by working the string from back to front and using a map."
] | {
"canonical_parameter_names": [
"s"
],
"leetcode_dataset_entry_point": "Solution().romanToInt",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:35:58",
"tests_total": 94,
"tests_dropped"... | cpp | {
"raw_signature": "int romanToInt(string s) {",
"container": "Solution",
"callable": "romanToInt",
"parameters": [
{
"name": "s",
"type": "string"
}
],
"return_type": "int",
"signature_source": "leetcode/codeSnippets",
"type_source": "leetcode/codeSnippets"
} |
14 | longest-common-prefix | Easy | Write a function to find the longest common prefix string amongst an array of strings.
If there is no common prefix, return an empty string "".
Example 1:
Input: strs = ["flower","flow","flight"]
Output: "fl"
Example 2:
Input: strs = ["dog","racecar","car"]
Output: ""
Explanation: There is no common prefix among t... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(strs = ['hello', 'helium', 'helper']) == \"hel\"\n assert candidate(strs = ['a']) == \"a\"\n assert candidate(strs = ['', '', '', '']) == \"\"\n assert candidate(strs = ['apple', 'app', 'apricot']) == \"ap\"\n assert candidat... | [
"[\"flower\",\"flow\",\"flight\"]",
"[\"dog\",\"racecar\",\"car\"]"
] | [] | {
"canonical_parameter_names": [
"strs"
],
"leetcode_dataset_entry_point": "Solution().longestCommonPrefix",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:36:00",
"tests_total": 148,
"t... | cpp | {
"raw_signature": "string longestCommonPrefix(vector<string>& strs) {",
"container": "Solution",
"callable": "longestCommonPrefix",
"parameters": [
{
"name": "strs",
"type": "vector<string>&"
}
],
"return_type": "string",
"signature_source": "leetcode/codeSnippets",
"type_source": "... |
15 | 3sum | Medium | Given an integer array nums, return all the triplets [nums[i], nums[j], nums[k]] such that i != j, i != k, and j != k, and nums[i] + nums[j] + nums[k] == 0.
Notice that the solution set must not contain duplicate triplets.
Example 1:
Input: nums = [-1,0,1,2,-1,-4]
Output: [[-1,-1,2],[-1,0,1]]
Explanation:
nums[0] + ... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(nums = [-2, 0, 0, 2, 2]) == [[-2, 0, 2]]\n assert candidate(nums = [0, 0, 0]) == [[0, 0, 0]]\n assert candidate(nums = [-1, 0, 1, 2, -1, -4]) == [[-1, -1, 2], [-1, 0, 1]]\n assert candidate(nums = [-2, 0, 1, 1, 2]) == [[-2, 0, 2... | [
"[-1,0,1,2,-1,-4]",
"[0,1,1]",
"[0,0,0]"
] | [
"So, we essentially need to find three numbers x, y, and z such that they add up to the given value. If we fix one of the numbers say x, we are left with the two-sum problem at hand!",
"For the two-sum problem, if we fix one of the numbers, say x, we have to scan the entire array to find the next number y, which ... | {
"canonical_parameter_names": [
"nums"
],
"leetcode_dataset_entry_point": "Solution().threeSum",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:36:03",
"tests_total": 104,
"tests_droppe... | cpp | {
"raw_signature": "vector<vector<int>> threeSum(vector<int>& nums) {",
"container": "Solution",
"callable": "threeSum",
"parameters": [
{
"name": "nums",
"type": "vector<int>&"
}
],
"return_type": "vector<vector<int>>",
"signature_source": "leetcode/codeSnippets",
"type_source": "le... |
16 | 3sum-closest | Medium | You are given an integer array nums of length n and an integer target.
Find three integers at distinct indices in nums such that the sum is closest to target.
Return the sum of the three integers.
You may assume that each input would have exactly one solution.
Example 1:
Input: nums = [-1,2,1,-4], target = 1
Outpu... | {
"language": "python",
"source": "def check(candidate):\n assert candidate(nums = [1, 2, 4, 8, 16, 32, 64, 128],target = 82) == 82\n assert candidate(nums = [1, 1, 1, 0],target = -100) == 2\n assert candidate(nums = [-10, -2, -5, -1],target = -12) == -13\n assert candidate(nums = [-5, -4, -3, -2, -1],t... | [
"[-1,2,1,-4]\n1",
"[0,0,0]\n1"
] | [] | {
"canonical_parameter_names": [
"nums",
"target"
],
"leetcode_dataset_entry_point": "Solution().threeSumClosest",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:36:05",
"tests_total":... | cpp | {
"raw_signature": "int threeSumClosest(vector<int>& nums, int target) {",
"container": "Solution",
"callable": "threeSumClosest",
"parameters": [
{
"name": "nums",
"type": "vector<int>&"
},
{
"name": "target",
"type": "int"
}
],
"return_type": "int",
"signature_sou... |
17 | letter-combinations-of-a-phone-number | Medium | Given a string containing digits from 2-9 inclusive, return all possible letter combinations that the number could represent. Return the answer in any order.
A mapping of digits to letters (just like on the telephone buttons) is given below. Note that 1 does not map to any letters.
Example 1:
Input: digits = "23"
Ou... | {
"language": "python",
"source": "def answers_match(result, expected):\n \"\"\"The answer may list its items in any order (only the top-level order is free).\"\"\"\n if isinstance(result, list) and isinstance(expected, list):\n return sorted(result, key=repr) == sorted(expected, key=repr)\n return ... | [
"\"23\"",
"\"2\""
] | [] | {
"canonical_parameter_names": [
"digits"
],
"leetcode_dataset_entry_point": "Solution().letterCombinations",
"problem_source": "newfacade/LeetCodeDataset",
"description_source": "leetcode.com",
"interface_source": "leetcode/codeSnippets",
"fetched_at": "2026-09-30T19:36:07",
"tests_total": 79,
"t... | cpp | {
"raw_signature": "vector<string> letterCombinations(string digits) {",
"container": "Solution",
"callable": "letterCombinations",
"parameters": [
{
"name": "digits",
"type": "string"
}
],
"return_type": "vector<string>",
"signature_source": "leetcode/codeSnippets",
"type_source": "... |
LeetCode multilingual benchmark dataset
LeetCode problems for the msl-multilingual-self-learning benchmark, flattened to
one row per (problem, language) in 9 languages, stored as data/<split>/<lang>-NNN.jsonl.
Each row has the interface for its language (from LeetCode's code snippets),
the shared canonical_tests (Python asserts from newfacade/LeetCodeDataset),
the problem_description (from the LeetCode page) and metadata.
Tests that break the problem's Constraints, do not fit a declared type in some
language, or expect inf/nan were removed; problems that take trees or linked
lists, modify their input in place, or kept fewer than 10 tests were dropped.
The test split also drops problems with several valid answers or whose text
refers to a figure. Where answers may come in any order or are decimals,
the asserts call answers_match (defined at the top of the test source), and
canonical_tests.comparison says which rule applies ("unordered", "float" or "exact").
public_tests holds the inputs of LeetCode's Testcase panel (one string per
case, one JSON value per line in parameter order; LeetCode publishes no
outputs), and hints the page's hints as plain text.
reports/dropped.md lists every dropped problem and test count, and
reports/<split>.json has the details.
Splits
train: 2103 problems x 9 languages = 18927 rows (2641 in the source dataset)test: 191 problems x 9 languages = 1719 rows (228 in the source dataset)
Load
from datasets import load_dataset
train = load_dataset("neulab/leetcode", split="train")
test = load_dataset("neulab/leetcode", split="test")
python_rows = test.filter(lambda r: r["language"] == "python")
- Downloads last month
- 194