Showing posts with label Reservoir Sampling. Show all posts
Showing posts with label Reservoir Sampling. Show all posts

LeetCode 398 - Random Pick Index


http://bookshadow.com/weblog/2016/09/11/leetcode-random-pick-index/
Given an array of integers with possible duplicates, randomly output the index of a given target number. You can assume that the given target number must exist in the array.
Note:
The array size can be very large. Solution that uses too much extra space will not pass the judge.
Example:
int[] nums = new int[] {1,2,3,3,3};
Solution solution = new Solution(nums);

// pick(3) should return either index 2, 3, or 4 randomly. Each index should have equal probability of returning.
solution.pick(3);

// pick(1) should return 0. Since in the array only nums[0] is equal to 1.
solution.pick(1);
https://discuss.leetcode.com/topic/58322/what-on-earth-is-meant-by-too-much-memory
  1. Like mine, O(N) memory, O(N) init, O(1) pick.
  2. Like @dettier's Reservoir Sampling. O(1) init, O(1) memory, but O(N) to pick.
  3. Like @chin-heng's binary search: O(N) memory, O(N lg N) init, O(lg N) pick.
X. O(n) init, O(1) pick
    public Solution(int[] nums) {
        for (int i=0; i<nums.length; i++) {
            int num = nums[i];
            if (!indexes.containsKey(num))
                indexes.put(num, new ArrayList<Integer>());
            indexes.get(num).add(i);
        }
    }
    
    public int pick(int target) {
        List<Integer> indexes = this.indexes.get(target);
        int i = (int) (Math.random() * indexes.size());
        return indexes.get(i);
    }
    
    private Map<Integer, List<Integer>> indexes = new HashMap<>();
X. O(nlogn) init, O(logn) pick
https://discuss.leetcode.com/topic/58295/share-my-c-solution-o-lg-n-to-pick-o-nlg-n-for-sorting
Pre-process sorting for O(nlg(n))
Pick in O(lg(n)) using binary search
O(n) space to store value/index pairs
    typedef pair<int, int> pp; // <value, index>

    static bool comp(const pp& i, const pp& j) { return (i.first < j.first); }

    vector<pp> mNums;

    Solution(vector<int> nums) {
        for(int i = 0; i < nums.size(); i++) {
            mNums.push_back(pp({nums[i], i}));
        }
        sort(mNums.begin(), mNums.end(), comp);
    }

    int pick(int target) {
        pair<vector<pp>::iterator, vector<pp>::iterator> bounds = equal_range(mNums.begin(), mNums.end(), pp({target,0}), comp);
        int s = bounds.first - mNums.begin();
        int e = bounds.second - mNums.begin();
        int r = e - s;
        return mNums[s + (rand() % r)].second;
    }

X. O(1) init, O(n) pick
https://discuss.leetcode.com/topic/58301/simple-reservoir-sampling-solution
So basically you may return -1 at the end? Correct me if I am wrong, I think the question requires a result which has been guaranteed.
Oops, I got it. Actually random.nextInt(++counter) this will guarantee that at least the first one will be picked u
 this is because we need to save current element with probability 1 / (total # of encountered candidates). So it's 100% for first target element, 50% for the second one, and so on...
To those who don't understand why it works. Consider the example in the OJ
{1,2,3,3,3} with target 3, you want to select 2,3,4 with a probability of 1/3 each.
2 : It's probability of selection is 1 * (1/2) * (2/3) = 1/3
3 : It's probability of selection is (1/2) * (2/3) = 1/3
4 : It's probability of selection is just 1/3
So they are each randomly selected.
In rand.nextInt(n)n sure is the exclusive boundary. But for n <= 0, the method will throw an IllegalArgumentExceptionnmust be a positive number.
So counter++ is wrong.
r.nextInt(1) always return 0
public class Solution {

    int[] nums;
    Random rnd;

    public Solution(int[] nums) {
        this.nums = nums;
        this.rnd = new Random();
    }
    
    public int pick(int target) {
        int result = -1;
        int count = 0;
        for (int i = 0; i < nums.length; i++) {
            if (nums[i] != target)
                continue;
            if (rnd.nextInt(++count) == 0)
                result = i;
        }
        
        return result;
    }
}


https://discuss.leetcode.com/topic/58467/java-o-n-variant-of-reservoir-sampling
https://discuss.leetcode.com/topic/58356/o-n-for-java-any-other-good-idea
    private int[] nums;

    public Solution(int[] nums) {
        this.nums = nums;
    }
    
    private Random r = new Random();
    
    public int pick(int target) {
        int ret = -1;
        if (nums == null) {
            return ret;
        }
        int upbound = 1;
        for (int i = 0; i < nums.length; i++) {
            if (nums[i] == target) {
                if (r.nextInt(upbound) == 0) {
                    ret = i;
                } 
                upbound++;
            }
        }
        return ret;
    }

def __init__(self, nums): """ :type nums: List[int] :type numsSize: int """ size = len(nums) self.next = [0] * (size + 1) self.head = collections.defaultdict(int) for i, n in enumerate(nums): self.next[i + 1] = self.head[n] self.head[n] = i + 1 def pick(self, target): """ :type target: int :rtype: int """ cnt = 0 idx = self.head[target] while idx > 0: cnt += 1 idx = self.next[idx] c = int(random.random() * cnt) idx = self.head[target] for x in range(c): idx = self.next[idx] return idx - 1

X. Offline
https://discuss.leetcode.com/topic/58394/simple-java-solution-o-n
This is appropriate when this api is called many times
    static int[] nums; 

    public Solution(int[] nums) {
        this.nums=nums;
    }
    
    public int pick(int target) {
        
        List<Integer> mm = new ArrayList<Integer>();
     
     for(int i=0; i<nums.length; i++){
      if(nums[i]==target)
       mm.add(i);
     }
     Random rn=new Random();
     return mm.get(rn.nextInt(mm.size()));
    }
https://tenderleo.gitbooks.io/leetcode-solutions-/content/GoogleMedium/398.html
https://github.com/mintycc/OnlineJudge-Solutions/blob/master/Leetcode/398_Random_Pick_Index.java
1. reservoir sampling: n个数⾥里里⾯面随机选k个数,要求概率相等。
2. 变种:给⼀一⼀一个vector,求最⼤大元素的index。如果有多个最⼤大元素,
均匀地随机返回任意⼀一⼀一个index。⽐比⽐比如:[1, 2, 3, 3],随机返回2,3,
每个的概率是50%。解法:扫 ⼀一遍记录最⼤大值和index,返回的时候⽤用随机
数rand()⼀一⼀一下(只记录最⼤大值的所有index即可)


LeetCode 382 - Linked List Random Node


https://discuss.leetcode.com/topic/53738/o-n-time-o-1-space-java-solution
Given a singly linked list, return a random node's value from the linked list. Each node must have the same probability of being chosen.
Follow up:
What if the linked list is extremely large and its length is unknown to you? Could you solve this efficiently without using extra space?
Example:


// Init a singly linked list [1,2,3].
ListNode head = new ListNode(1);
head.next = new ListNode(2);
head.next.next = new ListNode(3);
Solution solution = new Solution(head);

// getRandom() should return either 1, 2, or 3 randomly. Each element should have equal probability of returning.
solution.getRandom();
https://leetcode.com/problems/linked-list-random-node/discuss/85662/Java-Solution-with-cases-explain
After I read this one: http://blog.jobbole.com/42550/, it comes with a simple example and I understood suddenly, and write the code by myself. I translate it to English, so more people can benefit from it.
Start...
When we read the first node head, if the stream ListNode stops here, we can just return the head.val. The possibility is 1/1.
When we read the second node, we can decide if we replace the result r or not. The possibility is 1/2. So we just generate a random number between 0 and 1, and check if it is equal to 1. If it is 1, replace r as the value of the current node, otherwise we don't touch r, so its value is still the value of head.
When we read the third node, now the result r is one of value in the head or second node. We just decide if we replace the value of r as the value of current node(third node). The possibility of replacing it is 1/3, namely the possibility of we don't touch r is 2/3. So we just generate a random number between 0 ~ 2, and if the result is 2 we replace r.


We can continue to do like this until the end of stream ListNode.

https://discuss.leetcode.com/topic/53740/reservoir-sampling-java-solution
    ListNode head;
    Random random;
    /** @param head The linked list's head. Note that the head is guanranteed to be not null, so it contains at least one node. */
    public Solution(ListNode head) {
        this.head = head;
        random = new Random();
    }
    
    /** Returns a random node's value. */
    public int getRandom() {
        ListNode result = head;
        ListNode cur = head;
        int size = 1;
        while (cur != null) {
            if (random.nextInt(size) == 0) {
                result = cur;
            }
            size++;
            cur = cur.next;
        }
        
        return result.val;
    }
https://discuss.leetcode.com/topic/53738/o-n-time-o-1-space-java-solution
    ListNode head = null;
    Random randomGenerator = null;
    public Solution(ListNode head) {
        this.head = head;
        this.randomGenerator = new Random();

    }
    
    /** Returns a random node's value. */
    public int getRandom() {
        ListNode result = null;
        ListNode current = head;
        
        for(int n = 1; current!=null; n++) {
            if (randomGenerator.nextInt(n) == 0) {
                result = current;
            }
            current = current.next;
        }
        
        return result.val;
        
    }

https://discuss.leetcode.com/topic/53739/straight-forward-java-solution

Random Maximum | tech::interview


Random Maximum | tech::interview
给你一个array,返回array里面最大数字的index,但是必须是最大数字里面随机的一个index。比如[2,1,2,1,5,4,5,5]必须返回[4,6,7]中的随机的一个数字,要求O(1)space。
出现过很多次的FB的题,naive的做法是先扫一遍,找出最大值和最大值的个数。然后从头再扫一遍即可。
用Reservoir Sampling思路做,one pass就可以了。

int random_max(const vector<int>& nums) {
int ret = 0, count = 0, max = INT_MIN;
for(int i = 0; i < nums.size(); ++i) {
if(nums[i] > max) {
max = nums[i];
count = 1;
ret = i;
} else if(nums[i] == max) {
if ((rand() % ++count) == 0) ret = i;
}
}
return ret;
}

Read full article from Random Maximum | tech::interview

Return Index of Maximum with Equal Possibility - Facebook


https://gist.github.com/gcrfelix/21c5eca538bec217f837
http://www.1point3acres.com/bbs/forum.php?mod=viewthread&tid=123655&highlight=facebook
randomly return the index of max element in array
给一个全是数字的数组,随机返回0到当前位置中最大值得坐标
比如【1,2,3,3,3,3,1,2】
在最后一个2的时候有4个3都是最大值,要按1/4的概率返回其中一个3的index

最简单的解法当然是跑一遍array,把max elements的index记录到list中,然后random一下list,即可满足题意。

这样是O(N) time, O(N) space

follow up: 如果要求O(1)space呢?(或者说如果是个infinite array呢)

这时候不能记录 index 了,只能在遍历数组时,记录一个当前遇到的最大值:max,以及最大值的个数:counter
同时维持一个 ret 变量,每次遇到 max 时,以 1/count 的概率更新 ret 为当前位置。
.鐣欏璁哄潧-涓€浜�-涓夊垎鍦�
这样最后能保证 ret 是随机的一个max的index么,可以的,比如:

【1,2,3,3,3,3,1,2】
. visit 1point3acres.com for more.
可以用递推证明,比如遇到第k个max时,假设它被返回的概率是随机的,即1/k,那么第k + 1个max出现的时候,按我们的替换方式,有 k/k+1的可能k + 1不会被选中,也就是第k个max 幸存的概率为 被选中的概率*不被k+1换走的概率 = 1/k * k/(k + 1) = 1/(k + 1)
public int findMax(int[] array) {
  int len = array.length;
  int result = -1;
  int max = Integer.MIN_VALUE;
  int count = 0;
  for(int i=0; i<len; i++) {
    if(array[i] == max) {
      count ++;
      int num = new Random.nextInt(count);
      if(num == 0) {
        result = i;
      }
    } else if(max == Integer.MIN_VALUE || array[i] > max) {
      max = array[i];
      result = i;
      count = 1;
    }
  }
  return result;

Uniform Random sampling of Unbalanced Binary Tree


Uniform Random sampling of Unbalanced Binary Tree - Algorithms and Problem SolvingAlgorithms and Problem Solving
Given an unbalanced binary tree, write code to select k sample node at random
straightforward solution would be to list all the nodes in any traversal order (BFS, DFS, etc) and find random k index from the n nodes in the list. But how this approach would behave if n is very large and k is very small? or n is unknown? When n is very large the probability of choosing an element from n , k/n is very very small number when n>>k. Usual method for generating (e.g. random()*(i+1) or rand()%(i+1)). is very skewed. The distribution of selection probability trends to have long negative tail. So, obviously you can't guarantee an uniform sampling.

We need to find the samples for a binary tree. Without knowing the total number of elements we can find k random sample by keeping an index starting at index=0 and increment by one whenever we come across a tree node. Using such an index we can apply reservoir sampling while traversing the tree in any traversal order as follows in O(n) time –

public static TreeNode[] randomKSampleTreeNode(TreeNode root, int k){
 TreeNode[] reservoir = new TreeNode[k];
 Queue<TreeNode> queue = new LinkedList<TreeNode>();
 queue.offer(root);
 int index = 0;
 
 //copy first k elements into reservoir 
 while(!queue.isEmpty() && index < k){
  TreeNode node = queue.poll();
  reservoir[index++] = node;
  if(node.left != null){
   queue.offer(node.left);
  }
  if(node.right != null){
   queue.offer(node.right);
  }
 }
 
 //for index k+1 to the last node of the tree select random index from (0 to index) 
 //if random index is less than k than replace reservoir node at this index by 
 //current node
 while(!queue.isEmpty()){
  TreeNode node = queue.poll();
  int j = (int) Math.floor(Math.random()*(index+1));
  index++;
  
  if(j < k){
   reservoir[j] = node;
  }
 
  if(node.left != null){
   queue.offer(node.left);
  }
  if(node.right != null){
   queue.offer(node.right);
  }
 }
 
 return reservoir;
}
http://algobox.org/random-node-in-binary-tree/
Given an unbalanced binary tree, write code to select a node at random (each node has an equal probability of being selected).
The solutions depend on the actual requirements.

If the query happens only one time
1. Traversal the tree using Morris Inorder Traversal and count the total number of nodes n.
2. Get a random number i in [1, n].
3. Traversal the tree using Morris Inorder Traversal and return the ith node.
The algorithm is O(n) time with O(1) extra space.

If the query happens many times, but tree is immutable
For a immutable tree, we can preprocess to get a direct access table to all the nodes. In other words, we create an array nodes with each element point to one unique node in the tree.
Then the query becomes:
1. Get a random number i in [0, n)
2. Return nodes[i]
The query is then O(1).

If the query happens many times, but tree is mutable
We can preprocess the tree with argumentation of number of nodes in the sub tree. The query is then O(h).
Read full article from Uniform Random sampling of Unbalanced Binary Tree - Algorithms and Problem SolvingAlgorithms and Problem Solving

Labels

LeetCode (1432) GeeksforGeeks (1122) LeetCode - Review (1067) Review (882) Algorithm (668) to-do (609) Classic Algorithm (270) Google Interview (237) Classic Interview (222) Dynamic Programming (220) DP (186) Bit Algorithms (145) POJ (141) Math (137) Tree (132) LeetCode - Phone (129) EPI (122) Cracking Coding Interview (119) DFS (115) Difficult Algorithm (115) Lintcode (115) Different Solutions (110) Smart Algorithm (104) Binary Search (96) BFS (91) HackerRank (90) Binary Tree (86) Hard (79) Two Pointers (78) Stack (76) Company-Facebook (75) BST (72) Graph Algorithm (72) Time Complexity (69) Greedy Algorithm (68) Interval (63) Company - Google (62) Geometry Algorithm (61) Interview Corner (61) LeetCode - Extended (61) Union-Find (60) Trie (58) Advanced Data Structure (56) List (56) Priority Queue (53) Codility (52) ComProGuide (50) LeetCode Hard (50) Matrix (50) Bisection (48) Segment Tree (48) Sliding Window (48) USACO (46) Space Optimization (45) Company-Airbnb (41) Greedy (41) Mathematical Algorithm (41) Tree - Post-Order (41) ACM-ICPC (40) Algorithm Interview (40) Data Structure Design (40) Graph (40) Backtracking (39) Data Structure (39) Jobdu (39) Random (39) Codeforces (38) Knapsack (38) LeetCode - DP (38) Recursive Algorithm (38) String Algorithm (38) TopCoder (38) Sort (37) Introduction to Algorithms (36) Pre-Sort (36) Beauty of Programming (35) Must Known (34) Binary Search Tree (33) Follow Up (33) prismoskills (33) Palindrome (32) Permutation (31) Array (30) Google Code Jam (30) HDU (30) Array O(N) (29) Logic Thinking (29) Monotonic Stack (29) Puzzles (29) Code - Detail (27) Company-Zenefits (27) Microsoft 100 - July (27) Queue (27) Binary Indexed Trees (26) TreeMap (26) to-do-must (26) 1point3acres (25) GeeksQuiz (25) Merge Sort (25) Reverse Thinking (25) hihocoder (25) Company - LinkedIn (24) Hash (24) High Frequency (24) Summary (24) Divide and Conquer (23) Proof (23) Game Theory (22) Topological Sort (22) Lintcode - Review (21) Tree - Modification (21) Algorithm Game (20) CareerCup (20) Company - Twitter (20) DFS + Review (20) DP - Relation (20) Brain Teaser (19) DP - Tree (19) Left and Right Array (19) O(N) (19) Sweep Line (19) UVA (19) DP - Bit Masking (18) LeetCode - Thinking (18) KMP (17) LeetCode - TODO (17) Probabilities (17) Simulation (17) String Search (17) Codercareer (16) Company-Uber (16) Iterator (16) Number (16) O(1) Space (16) Shortest Path (16) itint5 (16) DFS+Cache (15) Dijkstra (15) Euclidean GCD (15) Heap (15) LeetCode - Hard (15) Majority (15) Number Theory (15) Rolling Hash (15) Tree Traversal (15) Brute Force (14) Bucket Sort (14) DP - Knapsack (14) DP - Probability (14) Difficult (14) Fast Power Algorithm (14) Pattern (14) Prefix Sum (14) TreeSet (14) Algorithm Videos (13) Amazon Interview (13) Basic Algorithm (13) Codechef (13) Combination (13) Computational Geometry (13) DP - Digit (13) LCA (13) LeetCode - DFS (13) Linked List (13) Long Increasing Sequence(LIS) (13) Math-Divisible (13) Reservoir Sampling (13) mitbbs (13) Algorithm - How To (12) Company - Microsoft (12) DP - Interval (12) DP - Multiple Relation (12) DP - Relation Optimization (12) LeetCode - Classic (12) Level Order Traversal (12) Prime (12) Pruning (12) Reconstruct Tree (12) Thinking (12) X Sum (12) AOJ (11) Bit Mask (11) Company-Snapchat (11) DP - Space Optimization (11) Dequeue (11) Graph DFS (11) MinMax (11) Miscs (11) Princeton (11) Quick Sort (11) Stack - Tree (11) 尺取法 (11) 挑战程序设计竞赛 (11) Coin Change (10) DFS+Backtracking (10) Facebook Hacker Cup (10) Fast Slow Pointers (10) HackerRank Easy (10) Interval Tree (10) Limited Range (10) Matrix - Traverse (10) Monotone Queue (10) SPOJ (10) Starting Point (10) States (10) Stock (10) Theory (10) Tutorialhorizon (10) Kadane - Extended (9) Mathblog (9) Max-Min Flow (9) Maze (9) Median (9) O(32N) (9) Quick Select (9) Stack Overflow (9) System Design (9) Tree - Conversion (9) Use XOR (9) Book Notes (8) Company-Amazon (8) DFS+BFS (8) DP - States (8) Expression (8) Longest Common Subsequence(LCS) (8) One Pass (8) Quadtrees (8) Traversal Once (8) Trie - Suffix (8) 穷竭搜索 (8) Algorithm Problem List (7) All Sub (7) Catalan Number (7) Cycle (7) DP - Cases (7) Facebook Interview (7) Fibonacci Numbers (7) Flood fill (7) Game Nim (7) Graph BFS (7) HackerRank Difficult (7) Hackerearth (7) Inversion (7) Kadane’s Algorithm (7) Manacher (7) Morris Traversal (7) Multiple Data Structures (7) Normalized Key (7) O(XN) (7) Radix Sort (7) Recursion (7) Sampling (7) Suffix Array (7) Tech-Queries (7) Tree - Serialization (7) Tree DP (7) Trie - Bit (7) 蓝桥杯 (7) Algorithm - Brain Teaser (6) BFS - Priority Queue (6) BFS - Unusual (6) Classic Data Structure Impl (6) DP - 2D (6) DP - Monotone Queue (6) DP - Unusual (6) DP-Space Optimization (6) Dutch Flag (6) How To (6) Interviewstreet (6) Knapsack - MultiplePack (6) Local MinMax (6) MST (6) Minimum Spanning Tree (6) Number - Reach (6) Parentheses (6) Pre-Sum (6) Probability (6) Programming Pearls (6) Rabin-Karp (6) Reverse (6) Scan from right (6) Schedule (6) Stream (6) Subset Sum (6) TSP (6) Xpost (6) n00tc0d3r (6) reddit (6) AI (5) Abbreviation (5) Anagram (5) Art Of Programming-July (5) Assumption (5) Bellman Ford (5) Big Data (5) Code - Solid (5) Code Kata (5) Codility-lessons (5) Coding (5) Company - WMware (5) Convex Hull (5) Crazyforcode (5) DFS - Multiple (5) DFS+DP (5) DP - Multi-Dimension (5) DP-Multiple Relation (5) Eulerian Cycle (5) Graph - Unusual (5) Graph Cycle (5) Hash Strategy (5) Immutability (5) Java (5) LogN (5) Manhattan Distance (5) Matrix Chain Multiplication (5) N Queens (5) Pre-Sort: Index (5) Quick Partition (5) Quora (5) Randomized Algorithms (5) Resources (5) Robot (5) SPFA(Shortest Path Faster Algorithm) (5) Shuffle (5) Sieve of Eratosthenes (5) Strongly Connected Components (5) Subarray Sum (5) Sudoku (5) Suffix Tree (5) Swap (5) Threaded (5) Tree - Creation (5) Warshall Floyd (5) Word Search (5) jiuzhang (5)

Popular Posts