2012年5月3日 星期四

[Python Std Library] Numeric and Math Modules : itertools — Iterators for efficient looping


翻譯自 這裡
Preface :
New in version 2.3.
這個模組提供一些 iterator 相關的應用. 接著我們來看看上面提供的函數 :
- itertools.chain(*iterables)
Make an iterator that returns elements from the first iterable until it is exhausted, then proceeds to the next iterable, until all of the iterables are exhausted. Used for treating consecutive sequences as a single sequence. Equivalent to :
  1. def chain(*iterables):  
  2.     # chain('ABC', 'DEF') --> A B C D E F  
  3.     for it in iterables:  
  4.         for element in it:  
  5.             yield element  
>>> iter = itertools.chain([1, 2, 3], ['a', 'b', 'c'])
>>> for i in iter:
... print('i={0}'.format(i))
...
i=1
i=2
i=3
i=a
i=b
i=c

- classmethod chain.from_iterable(iterable)
New in version 2.6.
Alternate constructor for chain(). Gets chained inputs from a single iterable argument that is evaluated lazily.

- itertools.combinations(iterable, r)
New in version 2.6.
Return r length subsequences of elements from the input iterable.

Combinations are emitted in lexicographic sort order. So, if the input iterable is sorted, the combination tuples will be produced in sorted order.

Elements are treated as unique based on their position, not on their value. So if the input elements are unique, there will be no repeat values in each combination. Equivalent to :
  1. def combinations(iterable, r):  
  2.     # combinations('ABCD', 2) --> AB AC AD BC BD CD  
  3.     # combinations(range(4), 3) --> 012 013 023 123  
  4.     pool = tuple(iterable)  
  5.     n = len(pool)  
  6.     if r > n:  
  7.         return  
  8.     indices = range(r)  
  9.     yield tuple(pool[i] for i in indices)  
  10.     while True:  
  11.         for i in reversed(range(r)):  
  12.             if indices[i] != i + n - r:  
  13.                 break  
  14.         else:  
  15.             return  
  16.         indices[i] += 1  
  17.         for j in range(i+1, r):  
  18.             indices[j] = indices[j-1] + 1  
  19.         yield tuple(pool[i] for i in indices)  
The code for combinations() can be also expressed as a subsequence of permutations() after filtering entries where the elements are not in sorted order (according to their position in the input pool) :
  1. def combinations(iterable, r):  
  2.     pool = tuple(iterable)  
  3.     n = len(pool)  
  4.     for indices in permutations(range(n), r):  
  5.         if sorted(indices) == list(indices):  
  6.             yield tuple(pool[i] for i in indices)  
The number of items returned is n! /( r! (n-r)!) when n is the size of iterable

- itertools.combinations_with_replacement(iterable, r)
New in version 2.7.
Return r length subsequences of elements from the input iterable allowing individual elements to be repeated more than once.
Combinations are emitted in lexicographic sort order. So, if the input iterable is sorted, the combination tuples will be produced in sorted order.
Elements are treated as unique based on their position, not on their value. So if the input elements are unique, the generated combinations will also be unique.
使用範例如下 :

- itertools.compress(data, selectors)
New in version 2.7.
Make an iterator that filters elements from data returning only those that have a corresponding element in selectors that evaluates to True. Stops when either the data or selectors iterables has been exhausted. Example as below :

- itertools.count(start=0, step=1)
Changed in version 2.7: added step argument and allowed non-integer arguments.
Make an iterator that returns evenly spaced values starting with start. Often used as an argument to imap() to generate consecutive data points. Also, used withizip() to add sequence numbers. Example :

- itertools.cycle(iterable)
Make an iterator returning elements from the iterable and saving a copy of each. When the iterable is exhausted, return elements from the saved copy. Repeats indefinitely. Example :

- itertools.dropwhile(predicate, iterable)
Make an iterator that drops elements from the iterable as long as the predicate is true; afterwards, returns every element. Note, the iterator does not produce any output until the predicate first becomes false, so it may have a lengthy start-up time. Example :

- itertools.groupby(iterable[, key])
New in version 2.4.
Make an iterator that returns consecutive keys and groups from the iterable. The key is a function computing a key value for each element. If not specified or is None,key defaults to an identity function and returns the element unchanged. Generally, the iterable needs to already be sorted on the same key function.

The returned group is itself an iterator that shares the underlying iterable with groupby(). Because the source is shared, when the groupby() object is advanced, the previous group is no longer visible. So, if that data is needed later, it should be stored as a list. Below is a usage example :

- itertools.filterfalse(predicate, iterable)
Make an iterator that filters elements from iterable returning only those for which the predicate is False. If predicate is None, return the items that are false. The usage example as below :

- itertools.islice(iterable[, start], stop[, step])
Changed in version 2.5: accept None values for default start and step.
Make an iterator that returns selected elements from the iterable. If start is non-zero, then elements from the iterable are skipped until start is reached. Afterward, elements are returned consecutively unless step is set higher than one which results in items being skipped. If stop is None, then iteration continues until the iterator is exhausted, if at all; otherwise, it stops at the specified position. Unlike regular slicing, islice() does not support negative values for start, stop, or step.

If start is None, then iteration starts at zero. If step is None, then the step defaults to one. Below is the usage example :

- itertools.zip_longest(*iterables[, fillvalue])
New in version 2.6.
Make an iterator that aggregates elements from each of the iterables. If the iterables are of uneven length, missing values are filled-in with fillvalue. Iteration continues until the longest iterable is exhausted. If not specified, fillvalue defaults to None. Usage example as below :

- itertools.permutations(iterable[, r])
New in version 2.6.
Return successive r length permutations of elements in the iterable.

If r is not specified or is None, then r defaults to the length of the iterable and all possible full-length permutations are generated.

Permutations are emitted in lexicographic sort order. So, if the input iterable is sorted, the permutation tuples will be produced in sorted order.

- itertools.product(*iterables[, repeat])
New in version 2.6.
Cartesian product of input iterables.

Equivalent to nested for-loops in a generator expression. For example, product(A, B) returns the same as ((x,y) for x in A for y in B).

To compute the product of an iterable with itself, specify the number of repetitions with the optional repeat keyword argument. For example, product(A, repeat=4)means the same as product(A, A, A, A). Below is the usage example :

- itertools.repeat(object[, times])
Make an iterator that returns object over and over again. Runs indefinitely unless the times argument is specified :

- itertools.starmap(function, iterable)
Changed in version 2.6: Previously, starmap() required the function arguments to be tuples. Now, any iterable is allowed.
Make an iterator that computes the function using arguments obtained from the iterable.

- itertools.takewhile(predicate, iterable)
Make an iterator that returns elements from the iterable as long as the predicate is true :

- itertools.tee(iterable[, n=2])
New in version 2.4.
Return n independent iterators from a single iterable.

Supplement :
* Recipes of itertools
This section shows recipes for creating an extended toolset using the existing itertools as building blocks...

2012年5月1日 星期二

[Python Std Library] Data Types : sched — Event scheduler


翻譯自 這裡
Preface :
sched 模組提供類別作為一般用途的時程事件功能. 首先來看看建構子 :
- class sched.scheduler(timefunc, delayfunc)
The scheduler class defines a generic interface to scheduling events. It needs two functions to actually deal with the “outside world” — timefunc should be callable without arguments, and return a number (the “time”, in any units whatsoever). The delayfunc function should be callable with one argument, compatible with the output of timefunc, and should delay that many time units. delayfunc will also be called with the argument 0 after each event is run to allow other threads an opportunity to run in multi-threaded applications.

接著我們來看一個簡單範例, 在 5秒, 10秒 後列印時間 :
>>> import sched, time
>>> s = sched.scheduler(time.time, time.sleep)
>>> def print_time(): print("\t[Info] From print_time: {0}".format(time.time()))
...
>>> def print_some_times():
... print("\t[Info] Now={0}...".format(time.time()))
... s.enter(5, 1, print_time, ()) # 添加五秒後, 執行 print_time() 事件.
... s.enter(10, 1, print_time, ()) # 添加十秒後, 執行 print_time() 事件.
... s.run() # 開始 Event schedule
... print("\t[Info] Now={0}...".format(time.time()))
...
>>> print_some_times()
[Info] Now=1333436870.441...
[Info] From print_time: 1333436875.442
[Info] From print_time: 1333436880.441
 # 與前一事件相差五秒
[Info] Now=1333436880.444...

事實上 sched 模組在處理多線程時會有 Thread-safe 的問題, 如果在需要考慮多線程使用情況下, 請使用 threading.Timer. 底下為簡單使用範例 :
>>> import time
>>> from threading import Timer
>>> def print_time():
... print("\t[Info] From print_time: {0}".format(time.time()))
...
>>> def print_some_times():
... print("\t[Info] Now : {0}".format(time.time()))
... Timer(5, print_time, ()).start() # 開始 Thread-1, 在 5 秒後呼叫 print_time()
... Timer(10, print_time, ()).start() # 開始 Thread-2, 在 10 秒後呼叫 print_time()
... time.sleep(11) # Sleep while time-delay events execute
... print("\t[Info] Now : {0}".format(time.time()))
...
>>> print_some_times()
[Info] Now : 1333437800.668
[Info] From print_time: 1333437805.669
[Info] From print_time: 1333437810.669
[Info] Now : 1333437811.669

Scheduler Objects :
scheduler 物件提供下面函數使用 :
- scheduler.enterabs(time, priority, action, argument)
Schedule a new event. The time argument should be a numeric type compatible with the return value of the timefunc function passed to the constructor. Events scheduled for the same time will be executed in the order of their priority.

Executing the event means executing action(*argument). argument must be a sequence holding the parameters for action.

Return value is an event which may be used for later cancellation of the event (see cancel()).

- scheduler.enter(delay, priority, action, argument)
Schedule an event for delay more time units. Other than the relative time, the other arguments, the effect and the return value are the same as those forenterabs().

- scheduler.cancel(event)
Remove the event from the queue. If event is not an event currently in the queue, this method will raise a ValueError.

- scheduler.empty()
Return true if the event queue is empty.

- scheduler.run()
Run all scheduled events. This function will wait (using the delayfunc() function passed to the constructor) for the next event, then execute it and so on until there are no more scheduled events.

Either action or delayfunc can raise an exception. In either case, the scheduler will maintain a consistent state and propagate the exception. If an exception is raised by action, the event will not be attempted in future calls to run().

If a sequence of events takes longer to run than the time available before the next event, the scheduler will simply fall behind. No events will be dropped; the calling code is responsible for canceling events which are no longer pertinent.

- scheduler.queue
New in version 2.6.
Read-only attribute returning a list of upcoming events in the order they will be run. Each event is shown as a named tuple with the following fields: time, priority, action, argument.

[ Eclipse Plugin ] EGit : 從 Github Import 專案到本地端


前言 :
這裡要介紹如何透過 Eclipse 的 plugin EGit 來從 Github import 專案到本地端.

Create Repository in Github :
在一開始你必須先在 Github 建立一個 Repository. 有關建立的過程說明可以參考 Create Repository at GitHub. 這裡假設你已經完成建立一個 Repository 且其 HTTP 的路徑為https://johnklee@github.com/johnklee/CRFPrac.git

安裝 EGit :
你可以直接使用 Eclipse (這裡使用的 Eclipse 版本 Version: 3.7.2) 的選單 : Help > Install New Software

(在出現視窗的 Work with 輸入 EGit install repository URL path. 可以參考這裡)

接著按照指示一步步便可完成安裝 EGit plugin 到 Eclipse. 而在使用 EGit 之前還需要一些簡單的設定, 可以參考這裡.

Import from Github :
接著我們要從 Github 上 import 專案 CRFPrac 到本地端. 詳細步驟請參考下面說明 :
Step1 : 請執行 Eclipse 選單 File > Import


Step2 : 再出現視窗選擇 Git > Projects from Git


Step3 : 在出現視窗選擇 URI 並點擊 "Next"


Step4 : 接著請到 Github 專案 CRFPrac, 選擇使用 HTTP 並複製其專案對應的 URL : https://johnklee@github.com/johnklee/CRFPrac.git


Step5 : 在 Step3 後出現的視窗填入 Step4 的 URL 後點擊 "Next"


Step6 : 在出現視窗點擊 "Next" (預設是 master 的 branch)


Step7 : 選擇 Local repository 的路徑後, 點擊 "Next"


Step8 : 選擇 "Import from existing project" 並點擊 "Next"

(這邊的 CRFPrac 專案有把 .project 等系統檔上傳到 Github, 所以可以使用選項 "Import from existing project"!)

Step9 : 在上一步後出現下面視窗, 點擊 "Finish" 完成專案 CRFPrac import!


Step10 : 最後你應該發現專案成功 import 到 Eclipse 的 Package Explorer


Supplement :
* EGit Document - User Guide
If you're new to Git or distributed version control systems generally, then you might want to read Git for Eclipse Users first. If you need more details and background read the Git Community Book or Git Pro...

* [Git Pro] Ch1 : Getting Started
在本章會透過一些版本控制的緣由說明, 來引入 Git 工具的使用. 包括為什麼你需要 Git, 如何取得並安裝與開始使用它...

* [Git Pro] Ch2 : Git Basics - Part 1
這邊介紹如何初始化 Git 的 repository 與基本的指令操作如 commit, status, add 與 設定那些檔案要 track 那些不要...

* [Git Pro] Ch2 : Git Basics - Part 2
這邊介紹命令讓你可以檢視之前提交的紀錄與如何對已經 commit 的檔案 undo...

[Git 常見問題] error: The following untracked working tree files would be overwritten by merge

  Source From  Here 方案1: // x -----删除忽略文件已经对 git 来说不识别的文件 // d -----删除未被添加到 git 的路径中的文件 // f -----强制运行 #   git clean -d -fx 方案2: 今天在服务器上  gi...